# Search and edit files

> Search file contents, edit files with patches that check the file's revision, and read a suspended workspace's files without waking it.

## What the file tools do

Besides reading, writing, listing and removing files, a workspace can:

- **search** file contents under a directory, like `grep -rn`, without starting a command;
- **patch** a file: replace exact text, or the whole content, atomically, optionally only if the file still has the
  **revision** (the SHA-256 of its content) your change is based on;
- answer **reads of a suspended workspace** from its disk, without resuming it.

| Client | Version | Search | Patch | Revision of a file |
| --- | --- | --- | --- | --- |
| TypeScript SDK | `@shardflux/sdk` **(0.9.0+)** | `cell.files.search()` | `cell.files.patch()` | `cell.files.stat(path, { revision: true })`, `cell.files.readWithInfo()` |
| Python SDK | `shardflux` **(0.5.0+)** | `ws.files.search()` | `ws.files.patch()` | `ws.files.stat(path, revision=True)`, `ws.files.read_with_info()` |
| CLI | `@shardflux/cli` **(0.5.0+)** | `shard files search` | `shard files patch` | `shard files stat --revision` |
| Agent tools and MCP server | `@shardflux/sdk` **(0.9.0+)**, `@shardflux/mcp` **(0.4.0+)** | `search_files` | `edit_file` | read by `edit_file` itself |
| HTTP (cell endpoint) | `/v1` | `POST .../files/search` | `POST .../files/patch` | `GET .../files/stat?revision=true`, `X-File-Revision` |

They work the same on [file-first workspaces](https://docs.shardflux.dev/concepts/file-first.md).

## Search file contents

```ts
const cell = workspace.cell();
const hits = await cell.files.search('/home/user/project', 'TODO', { include: ['*.py'], contextLines: 1 });
for (const m of hits.matches) console.log(`${m.path}:${m.line}:${m.column}: ${m.text}`);
if (hits.truncated) console.log('stopped early:', hits.stop_reason);
```

```python
hits = ws.files.search("/home/user/project", "TODO", include=["*.py"], context_lines=1)
for m in hits["matches"]:
    print(f"{m['path']}:{m['line']}:{m['column']}: {m['text']}")
```

```sh
shard files search customer-42/main /home/user/project TODO --include '*.py'
shard files search customer-42/main /home/user/project 'def \w+_test' --regex -i --json
```

- The pattern is literal text, or an RE2 regular expression with `regex` (`--regex`). `caseInsensitive` /
  `case_insensitive` (`-i`) ignores case.
- `path` is a directory, or one file. Matches come in path order, each with `path`, the 1-based `line`, the 1-based
  byte `column` of the first match on the line, and `text` (the line, up to 1,000 bytes). `contextLines` (0-5) adds
  `before` and `after` lines. The CLI prints `path:line:column: text`, and context lines as `path-line- text`.
- `include` and `exclude` take up to 32 gitignore-style globs each: `*.py` matches a name at any depth, `src/**/*.ts`
  a path relative to `path`, a leading `/` anchors to `path`, and a trailing `/` matches directories only (`build/`).
  By default `.git` and `node_modules` are skipped; an `exclude` you give replaces that default (`[]` searches
  everything).
- Symbolic links are not followed. Binary files (a NUL byte in the first 8 KiB), special files, `/proc` and `/sys`,
  and files larger than `maxFileBytes` / `max_file_bytes` (default 1 MiB, at most 64 MiB) are skipped. Each line is
  searched in its first 1 MiB.
- A search stops at `maxMatches` / `max_matches` (default 200, at most 5,000; `--max` in the CLI), after 10 seconds, or
  at 4 MiB of results. Then `truncated` is `true` and `stop_reason` is `max_matches`, `budget` or `max_bytes`.
  `files_scanned` counts the files searched.
- A search is read-only: the SDKs retry it like a `GET`.

## Edit a file with a patch

```ts
const path = '/home/user/project/app.py';
const { revision } = await cell.files.stat(path, { revision: true });
const patched = await cell.files.patch({
  path,
  edits: [{ oldText: 'DEBUG = True', newText: 'DEBUG = False' }],   // must occur exactly once (or replaceAll)
  expectedRevision: revision,                                        // refused if the file changed meanwhile
});
patched.revision;   // the file's new revision: the expectedRevision of your next patch
```

```python
path = "/home/user/project/app.py"
revision = ws.files.stat(path, revision=True)["revision"]
patched = ws.files.patch(
    path,
    edits=[{"old_text": "DEBUG = True", "new_text": "DEBUG = False"}],  # must occur exactly once
    expected_revision=revision,  # refused if the file changed meanwhile
)
patched["revision"]  # the next expected_revision
```

```sh
REV=$(shard files stat customer-42/main /home/user/project/app.py --revision --json | jq -r .revision)
shard files patch customer-42/main /home/user/project/app.py --old "DEBUG = True" --new "DEBUG = False" --expected-revision "$REV"
```

- A patch takes exactly one of **`edits`** or **`content`**. Each edit's old text must occur exactly once in the file,
  unless `replaceAll` / `replace_all` (`--replace-all`) replaces every occurrence. Edits apply in order to the file's
  UTF-8 text, all of them or none. `content` replaces the whole file, or creates it (`--content-file F`, or `-` for
  standard input).
- **`expectedRevision`** / `expected_revision` (`--expected-revision`) makes the patch apply only to that revision of
  the file. `absent` requires that the file does not exist yet.
- A patch is atomic and durable: it is acknowledged after the file and its directory are written to disk. An existing
  file keeps its mode and owner. A new file gets `mode` (default `0644`); `createParents` / `create_parents` creates
  missing directories.
- The SDKs send an `Idempotency-Key` with every patch, so a retried request is applied once.
- The result has the file's new `revision`, its `previous_revision`, `bytes_written`, `durable`, `replacements` (the
  number of replaced occurrences) and `file` (its metadata). The CLI prints the new revision.
- A request is at most 7 MiB and a file at most 64 MiB; write larger files with a normal write. A patch has at most
  100 edits. A patch through a symbolic link is refused.

| Refusal | `details` | Meaning |
| --- | --- | --- |
| `409 conflict`, reason `revision_mismatch` | `current_revision` | The file changed since the revision you gave (or exists, with `absent`). Nothing changed: read it again and redo the edit. |
| `422 validation_failed`, reason `edit_not_found` | `index` | The edit at that position (from 0) does not occur in the file. |
| `422 validation_failed`, reason `edit_ambiguous` | `index` | That edit's old text occurs more than once: include more surrounding text, or replace all. |
| `422 validation_failed`, reason `edit_not_text` | | The file is not UTF-8 text. Replace it with `content` instead. |
| `422 validation_failed`, reason `patch_invalid` | | Not exactly one of `edits` and `content`. |
| `404 not_found` | | No such file, and the patch has `edits` (only `content` creates a file). |
| `413 payload_too_large` | `max_bytes` | The request is over 7 MiB. |

## Revisions

A file's revision is the SHA-256 of its content, as 64 lowercase hex characters. It is what `expectedRevision` takes.

| Where | Returns the revision |
| --- | --- |
| `stat(path, { revision: true })`, `stat(path, revision=True)`, `shard files stat --revision`, `GET .../files/stat?revision=true` | For regular files up to 256 MiB. |
| `readWithInfo(path)` (TypeScript), `read_with_info(path)` (Python), the `X-File-Revision` header of a read | For regular files up to 16 MiB: `{ data, size, revision, servedFrom }`. A read continued over several requests has a revision only if every part had the same one. |
| Every write and patch result | `revision`, the file's revision after the change (for a replace, the same value as a write's `sha256`). |

`read()` and `readText()` / `read_text()` still return bytes and text. On a [file-first
workspace](https://docs.shardflux.dev/concepts/file-first.md) revisions are returned for files of any size.

## Read a suspended workspace without waking it

Reads of a suspended workspace are answered from its disk when a host still holds that disk: no resume, no compute, and
the files as they were when it was suspended. The workspace stays suspended.

| Client | Reads served from the disk | How you can tell |
| --- | --- | --- |
| TypeScript SDK **(0.9.0+)** | `read`, `readText`, `readWithInfo`, `stat`, `list`, `search` | `servedFrom: 'disk'` (`readWithInfo`), `served_from: 'disk'` (search) |
| Python SDK **(0.5.0+)** | `read`, `read_text`, `read_with_info`, `stat`, `list`, `search` | `served_from == "disk"` |
| CLI **(0.5.0+)** | `files read`, `ls`, `stat`, `search` | a note on stderr; `served_from` with `--json` |
| MCP server **(0.4.0+)** | `read_file`, `list_files`, `search_files` | |
| HTTP (cell endpoint) | `GET .../files`, `files/stat`, `files/list`, `POST .../files/search` | `X-Served-From: disk` |

- Every other call wakes the workspace as before: writes, patches, commands.
- When the disk cannot answer (it is no longer on a host, or the read is too large for it), the read is refused with
  `409 workspace_not_running` (`details.reason`: `offline_unavailable` or `offline_budget`). The SDKs, the CLI and the
  MCP server then wake the workspace and read again.
- A read that races the workspace's resume can fail with `503 dependency_unavailable`
  (`details.reason: offline_changed`); the SDKs retry it, and the running workspace answers.
- These reads are not tool activity and are not billed as compute.
- A running workspace that its host has [parked](https://docs.shardflux.dev/concepts/lifecycle.md#idle-running-workspaces-are-parked) may be read
  the same way, without waking it, so `served_from: 'disk'` can appear for a running workspace too.
- The API issues tool tokens for suspended workspaces for these reads, so a client that had no token before the
  suspend can read too.

## Agent tools: search_files and edit_file

`workspaceTools()` in the TypeScript SDK **(0.9.0+)** and the MCP server **(0.4.0+)** add two tools with the `files`
permission:

| Tool | Arguments (required in bold) | Returns |
| --- | --- | --- |
| `search_files` | **`path`**, **`pattern`**, `regex`, `case_insensitive`, `include`, `exclude`, `max_matches` (1-5,000), `context_lines` (0-5) | `matches`, `truncated`, `stop_reason`, `omitted_matches`, `files_scanned` |
| `edit_file` | **`path`**, **`edits`** (1-100 of `{old_text, new_text, replace_all}`), `expected_revision` | `path`, `revision`, `previous_revision`, `replacements`, `bytes_written` |

- `search_files` returns whole matches up to the tools' output limit (64 KiB by default) and counts the rest in
  `omitted_matches`.
- `edit_file` pins every edit to a revision. When the model gives no `expected_revision`, the tool reads the file's
  revision first, so a change made in between fails the edit (`revision_mismatch`) instead of being overwritten.
  The `revision` it returns is the `expected_revision` of the next edit of the same file.
- The Python SDK's `workspace_tools()` does not have these two tools yet. Use `ws.files.search()` and
  `ws.files.patch()`.

See [Give your agent workspace tools](https://docs.shardflux.dev/guides/agent-tools.md) for the rest of the tools.

## Hosts that do not have these calls yet

While the fleet is upgraded, a workspace can still run on a host without search or patches. Then search and patch are
refused with `409 conflict`, `details.reason: host_feature_unavailable` and `details.feature` (`file_search` or
`file_patch`), `retryable: false`, and reads and stats come without revisions. Use another way until the workspace
runs on an upgraded host: `grep -rn` in a command instead of a search, or a read and a write instead of a patch.
