# Lifecycle

> Open, suspend, resume, fork, reset and delete Shardflux workspaces: states, idle suspend and parking, waiting, and what survives each transition.

## States and operations

A workspace's `state` (the API's `observed_state`) is one of:

| State | Meaning |
| --- | --- |
| `creating`, `starting` | The first open or a restart is booting it. |
| `running` | It runs, and tool calls are served. |
| `suspending` | Its memory, processes and disk are being checkpointed. Tool calls wait. |
| `suspended` | The checkpoint is stored durably and the VM is gone. No compute; the disk still counts as storage. |
| `resuming` | The checkpoint is being restored. |
| `forking` | A fork's copy is being created (the new workspace's state). |
| `failed` | A start failed. Opening the key again starts it again. |
| `deleting`, `deleted` | It is being or has been deleted. |

A running workspace that nobody uses is [parked](#idle-running-workspaces-are-parked) by its host. That is not a state:
it stays `running`. A [file-first workspace](https://docs.shardflux.dev/concepts/file-first.md) is `running` from its first open until it is deleted.

Every change of state is an **operation** with a kind (`open`, `suspend`, `resume`, `fork`, `snapshot`, `reset`,
`delete`) and a state of its own: `queued`, `capacity_pending`, `running`, then `succeeded`, `failed` or `canceled`. A
workspace runs one lifecycle operation at a time; a second one is refused with `409 conflict`
(`details.reason: operation_in_progress`), except that opening or resuming joins a start already in progress.

## Requested or finished

Lifecycle calls return as soon as the change is **requested**, or, with `wait`, once it has **finished**:

| Call | Returns when | Returns |
| --- | --- | --- |
| `await workspace.suspend()` | The suspend is requested (usually `queued`; the workspace still runs) | The operation |
| `await workspace.suspend({ wait: true })` | The suspend has finished; `workspace.state` is `suspended` | The succeeded operation |
| `ws.suspend()` (Python) | The suspend is requested | The operation |
| `ws.suspend(wait=True)` (Python) | The suspend has finished; `ws.state` is `suspended` | The succeeded operation |
| `shard ws suspend <key>` | The suspend is requested | Prints the operation |
| `shard ws suspend <key> --wait` | The suspend has finished | Prints the operation |

The same applies to `resume`, `snapshot`, `delete`, `reset` and `close`. `open()` waits by default. `fork` waits by
default in Python and takes `{ wait: true }` in TypeScript.

## Open, create or reconnect

`open` is the one call you need for most of the lifecycle. It creates the workspace for a new key, returns a running
one, resumes a suspended one, and joins a start already in progress. It never resets anything. See
[What opening a key does](https://docs.shardflux.dev/concepts/workspaces.md#what-opening-a-key-does).

## Suspend a workspace

```ts
await workspace.suspend({ wait: true });   // memory and processes are checkpointed; compute stops
```

```python
ws.suspend(wait=True)
```

A suspend pauses the VM, stores its memory and disk durably, and then releases the CPU and memory. The workspace is
`suspended` only after the checkpoint is stored. If storing it fails, the suspend fails and the workspace runs again
from its local copy; nothing is silently lost.

Session workspaces cannot be suspended (`409`, `details.reason: session_lifetime`).

## Resume a workspace

You rarely need to resume explicitly:

- **Opening the key** resumes a suspended workspace.
- **A tool call wakes it.** An exec, file or other tool call on a suspended workspace resumes it (or joins the resume
  already running), then runs. A call made during a suspend or resume waits for it to finish. The call never runs
  twice. The SDKs bound the wait per call (default 120 seconds, `transitionTimeoutMs` in TypeScript,
  `transition_timeout` in Python); the CLI uses `--wake-timeout` and the MCP server `SHARDFLUX_WAKE_TIMEOUT_MS`.
- **The wake is one request** (`@shardflux/sdk` **0.9.0+**, `shardflux` **0.5.0+** for Python, CLI **0.5.0+**, MCP
  server **0.4.0+**). The client sends one resume that the API holds until the workspace runs, and the answer carries a
  tool token for that client, so the refused call is retried at once. An explicit `resume({ wait: true })`,
  `resume(wait=True)` or `shard ws resume --wait` is the same single request. Earlier versions wait for the resume
  operation and then fetch a token.
- **Reads do not wake it.** Reading, listing, stating and searching files of a suspended workspace are answered from
  its disk while a host still holds it, without a resume (`@shardflux/sdk` **0.9.0+**, `shardflux` **0.5.0+**, CLI
  **0.5.0+**, MCP server **0.4.0+**). See [Search and edit files](https://docs.shardflux.dev/guides/files.md#read-a-suspended-workspace-without-waking-it).
- Following an exec's output never wakes a workspace, so a suspend you asked for is respected.

```ts
await workspace.resume({ wait: true });           // explicit
await workspace.wake();                           // resume if needed; resolves once it runs
const cell = workspace.cell({ wake: null });      // opt out: calls fail with workspace_not_running instead
```

```python
ws.resume(wait=True)
ws.wake()                   # True if it resumed or waited, False if it was already running
cell = ws.cell(wake=None)   # opt out
```

A resume is admitted like a new start, so plan limits apply (see [Pricing and limits](https://docs.shardflux.dev/limits.md#what-happens-at-a-limit)).

## Fork a workspace

A fork creates a new workspace under a new key from the current state of a running or suspended one: a copy of its
disk and memory, restored into an independent VM with its own network identity. The source keeps running, or stays
suspended.

```ts
const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });
```

```python
copy = ws.fork("customer-42/experiment")   # waits until the copy runs
```

```sh
shard ws fork customer-42/main customer-42/experiment --wait
```

- The new key must be unused, including by deleted workspaces (`409`, `details.reason: key_in_use`).
- The fork keeps the source's template version, secret bindings and text inputs. It gets the source's caps unless you
  pass `caps`.
- A fork is persistent unless you pass `lifetime: 'session'`. Forking is how you keep a session workspace's state.
- A fork is a new start: plan limits apply to it as to an open.

## Snapshot a workspace

`workspace.snapshot({ label })` (TypeScript), `ws.snapshot(label=...)` (Python) records a checkpoint of a running or
suspended workspace without stopping it. While a running workspace is captured, reads work and writes wait. The API
has no call to list or restore snapshots yet; to branch from a state, [fork](#fork-a-workspace).

## Reset a workspace

`reset` wipes every change a workspace made and restarts it on its template version. It keeps the key, id, template
version, caps, secret bindings and volume attachments. Running processes end. It works only on layered workspaces
(`workspace.diskLayout === 'layered'`; others get `409`, `details.reason: legacy_disk_layout`). The previous state is
kept as a recovery point for 7 days.

```ts
await workspace.reset({ wait: true });
```

```sh
shard ws reset customer-42/main --yes --wait
```

A suspended workspace stays suspended and starts blank on its next resume.

## Delete a workspace

```ts
await workspace.delete({ wait: true });
```

```sh
shard ws delete customer-42/main --yes
```

Tool access ends at once and the workspace's storage is cleaned up. Deletion cannot be undone, and a persistent
workspace's key is never reused. The MCP server does not offer deletion to agents; use the SDK, the CLI or the
console.

A session workspace ends with `close()` (`shard ws close <key>`), which deletes it; its key then opens a new
workspace.

## Automatic suspend when idle

A persistent workspace suspends itself when it has been idle for its idle timeout. Nothing is killed: the suspend keeps
memory and processes like any other. A running workspace uses RAM GiB-hours for every second it is awake, idle or not,
so the shorter the idle tail after its last work, the less it costs you.

**What keeps a workspace awake:**

- a tool call (exec, files, terminal input, processes, git, browser);
- an attached exec output stream or terminal;
- a command started through exec, until it ends (for at most 1 hour, or its own timeout);
- a keepalive (`POST /v1/workspaces/{id}/keepalive` on the workspace's cell endpoint, for up to 12 hours).

A detached process, a dev server nobody is attached to, or CPU use alone does not keep it awake. Such processes
continue after the resume.

**Idle policy.** The default policy is `adaptive`: each workspace learns its own timeout, from 10 seconds to 4 hours.

- It learns from the workspace's idle periods, every pause of 5 seconds or more between one activity and the next.
  Pauses inside its usual active hours and outside them are learned separately.
- Short pauses, such as an agent thinking between tool calls, are not worth a suspend and a resume, so the timeout
  outlasts them. When a workspace is woken soon after an automatic suspend, its next timeouts get longer (a suspend
  you asked for never has this effect).
- Until the workspace has 8 idle periods of its own, its history is blended with that of its template, else its
  organization, else all workspaces. With no history at all, the timeout is 5 minutes.
- The work signals above always win: the policy only decides how long an idle workspace waits.

The other policies are `never` and `fixed:<seconds>` (60 to 604800). The workspace view reports under `idle` the
policy, the current timeout (`timeout_seconds`), what it is based on (`basis`, for example `learned from 37 idle
periods (active-hours)`, `template prior` or `default`), when it would suspend (`next_eligible_at`) and a pending
[suspend when idle](#suspend-when-idle) request (`suspend_request`).

Set it with the HTTP API; the SDKs have no dedicated method yet, so use their `request()` passthrough:

```ts
await cloud.request('PUT', `/v1/workspaces/${workspace.id}/idle-policy`, { json: { idle_policy: 'fixed:600' } });
```

```python
sf.request("PUT", f"/v1/workspaces/{ws.id}/idle-policy", json={"idle_policy": "never"})
```

`null` clears the workspace's own policy, so the template default (else `adaptive`) applies. Session workspaces have
no idle policy; they end after their idle timeout (10 minutes unless the template sets one).

### Suspend when idle

Your code knows when an agent's turn ends. Instead of waiting for the idle timeout, ask for a suspend once the
workspace has been idle for a short time. The request is stored with the workspace, so your process does not have to
stay around for the suspend.

```ts
await workspace.suspendWhenIdle({ afterSeconds: 60 });   // 30 to 3600
workspace.suspendRequest;                                // { requested_at, after_seconds, not_before }, or null
await workspace.cancelSuspendWhenIdle();                 // idempotent
```

```python
ws.suspend_when_idle(after_seconds=60)  # 30 to 3600
ws.suspend_request                      # SuspendRequest(requested_at, after_seconds, not_before), or None
ws.cancel_suspend_when_idle()           # idempotent
```

```sh
shard ws suspend customer-42/main --when-idle 1m      # 30s to 1h
shard ws suspend customer-42/main --cancel-when-idle
```

- The workspace is suspended once it has been idle for `after_seconds`, counted from the later of its last work and
  the request. `not_before` is the earliest suspend.
- A command still running, an attached exec output stream or terminal, or a keepalive postpones the suspend until
  `after_seconds` after it ends. The request stays.
- The next tool call on the workspace (the next turn) or a resume cancels the request. Asking again replaces it.
- It applies under every idle policy, `never` included, and never delays a suspend the policy would do sooner.
- If a suspend is already in progress, the call returns that operation and records nothing.
- A session workspace is refused (`409`, `details.reason: session_lifetime`), and so is a workspace that is not
  running (`not_running`).

It is in `@shardflux/sdk` **0.10.0+** (`cloud.workspaces.suspendWhenIdle(id, { afterSeconds })` by id), `shardflux`
**0.6.0+** on PyPI (`sf.workspaces.suspend_when_idle(workspace_id, after_seconds=60)`), `@shardflux/cli` **0.5.1+** and
the MCP server **0.4.1+** (`workspace_suspend` with `after_seconds`). Over HTTP it is
`POST /v1/workspaces/{id}/suspend-when-idle` with `{"after_seconds": 60}`, and `DELETE` on the same path cancels it.
The workspace view shows a pending request as `idle.suspend_request`. [Give your agent workspace
tools](https://docs.shardflux.dev/guides/agent-tools.md#suspend-when-the-turn-ends) shows the call at the end of an agent loop.

## Idle running workspaces are parked

Long before its idle timeout, a running workspace that nobody is using is **parked** by its host: after a few seconds
without activity its VM is paused and most of its memory is compressed, and after a longer pause the VM is saved to the
host's local disk. Parking is not a lifecycle state and needs nothing from you. The workspace stays `running`, keeps
its files, memory and processes, and the API answers as before. The next tool call wakes it first and then runs: about
a millisecond after a short pause, about a tenth of a second after a longer one.

- **What keeps it resident:** a tool call in progress, an attached exec output stream or terminal, a command started
  through exec until it ends, a keepalive, and processes that keep using CPU or the network.
- **Background processes that wait** (an idle dev server, a file watcher, a shell) are paused with the workspace and
  continue at the next wake. Timers inside the workspace fire late, never early, and the clock is right after the
  wake.
- **Network traffic does not wake it.** A process that only waits for data from the network is paused too, and a
  remote peer that expects a prompt answer may time out. A keepalive (`POST /v1/workspaces/{id}/keepalive` on the
  cell endpoint) keeps the workspace resident, and wakes it if it is parked.
- **Billing does not change.** A parked workspace is running: it uses RAM GiB-hours at its full memory allocation and
  counts toward running workspaces at once. CPU hours count the CPU time its processes use, which is next to none
  while it is parked. To stop RAM GiB-hours, suspend the workspace, or let the idle policy suspend it: parking neither
  delays nor replaces the automatic suspend.
- **Reads may not wake it.** Reading, listing, stating and searching files of a parked workspace can be answered from
  its disk (`X-Served-From: disk`) without waking it.
- **A wake can be refused for a moment.** When the host has no room to restore the workspace right away, a tool call
  is refused with `503 service_unavailable` (`details.reason: host_capacity`) and `Retry-After`; when a restore
  fails, with `wake_failed` (the saved state is intact). Both are retryable, and nothing was executed. The SDKs retry
  reads, searches and calls with an `Idempotency-Key`; other calls surface the error with `retryable: true`.

### Wake hint

A wake hint tells the host that a tool call is coming, so a parked workspace starts waking while your model is still
writing the call:

```ts
void workspace.hint().catch(() => {});   // @shardflux/sdk 0.9.0+: returns at once
```

```python
ws.hint()  # shardflux 0.5.0+: returns at once
```

- A hint is cheap and never waits. For a suspended workspace it starts the resume in the background (TypeScript
  `result.wake`, a promise; Python `WakeHint.wake`, a `Future`); `hint({ wake: null })` / `hint(wake=False)` only
  reports.
- The TypeScript agent tools send a hint when each tool call starts, except for `read_file`, `list_files` and
  `search_files`, which a sleeping workspace answers from its disk. `workspaceTools(ws, { hint: false })` turns that
  off. The MCP server **(0.4.0+)** does the same.
- Over HTTP, `POST /v1/workspaces/{id}/wake-hint` on the cell endpoint answers `202` with the `residency` the host
  found: `resident`, `frozen` (paused), `hibernated` (saved to the host's disk) or `restoring`. A suspended workspace
  answers `409 workspace_not_running`: resume it through the API. A hint that no tool call follows within 60 seconds
  is dropped. A hint is not tool activity and does not keep the workspace from its idle suspend.

## What survives each transition

Files and disk, and memory and running processes, are listed separately.

| Transition | Files and disk | Memory and running processes | Also |
| --- | --- | --- | --- |
| Open the same key again | Kept | Kept | Nothing is reset. |
| Parked while idle, then woken by a tool call | Kept | Kept: processes are paused and continue | The workspace stays `running` and is billed as running. Timers fire late. |
| Suspend (by you, when idle, or at a compute allowance or spend cap), then resume | Kept: every file write acknowledged before the suspend began, and installed packages | Kept: every process with its PID, memory, open files, working directory and environment, including detached processes, dev servers and exec or terminal sessions; listening sockets and loopback connections | Your connections to exec output and terminals end; reattach by offset (the SDK does). Remote peers may close connections while the workspace sleeps. |
| Resume after a platform runtime change (rare, see below) | Kept: as of the suspend | Ended: the workspace boots from its saved disk, as after a reboot | The resume result says `resume_path: "cold_boot"`, `memory_restored: false`, `cold_boot_reason: "runtime_changed"`. Start dev servers and background jobs again. |
| Fork | Copied into the new workspace | Copied into the new workspace, as an independent VM with a new network identity | The source is unchanged. Shared volumes are attached, not copied: both see the same data. |
| Snapshot | Unchanged | Unchanged | Writes wait during the capture. |
| Reset (layered workspaces) | Wiped back to the template | Ended | Key, id, template version, caps, secrets and volumes are kept; the old state is a recovery point for 7 days. |
| Session ends (`close()` or idle timeout) | Deleted | Ended | The key opens a new, empty workspace. |
| Delete | Deleted | Ended | The key of a persistent workspace is never reused. |
| A start fails with `capacity_unavailable` | Unchanged | Unchanged | Nothing was started; a suspended workspace stays suspended. |

A resume restores memory and disk from one checkpoint. Shared volumes are outside checkpoints: a suspend keeps the
attachment, and the resume mounts the volume's current contents.

### When a resume boots instead

A memory checkpoint can only be restored on a host that runs the same VM runtime (kernel, CPU family and Firecracker
snapshot format) as the host that saved it. We keep that runtime stable and roll hosts so that suspended workspaces
resume with their memory. When a runtime change is unavoidable (a security update, for example) and no host can
restore a workspace's memory any more, its next resume boots it from its saved disk instead of waiting forever:

- its files are exactly as they were at the suspend;
- its memory and processes are not restored: every process starts fresh, as after a reboot, and the template's
  startup runs its boot steps;
- the resume operation's result says so: `resume_path: "cold_boot"`, `memory_restored: false` and
  `cold_boot_reason: "runtime_changed"`. Every resume result has `memory_restored` (`true` when memory came back).

The clients surface it: `workspace.lastTiming.server.memoryRestored` and `coldBootReason` in TypeScript
(`@shardflux/sdk` **0.11.0+**), `memory_restored` and `cold_boot_reason` in Python (`shardflux` **0.7.0+**), a note
on stderr from the CLI (**0.6.0+**) and a notice the MCP server (**0.5.0+**) hands the agent, so it can restart what it
was running.

## Wait for an operation to finish

`wait` takes options: `{ timeoutMs, signal, onProgress }` in TypeScript (default 5 minutes), `timeout=` in seconds in
Python (default 300). When the time runs out, the SDK throws `OperationTimeoutError`, but the operation continues on
the server. Wait for it again:

```ts
import { OperationFailedError, OperationTimeoutError } from '@shardflux/sdk';

try {
  await workspace.resume({ wait: { timeoutMs: 60_000 } });
} catch (err) {
  if (err instanceof OperationTimeoutError) {
    await cloud.workspaces.waitForOperation(err.operationId);          // keep waiting
  } else if (err instanceof OperationFailedError && err.retryable) {
    // err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
  } else {
    throw err;
  }
}
```

```python
from shardflux import OperationFailedError, OperationTimeoutError

try:
    ws.resume(wait=True, timeout=60)
except OperationTimeoutError as err:
    sf.workspaces.wait_for_operation(err.operation_id)
except OperationFailedError as err:
    if not err.retryable:
        raise
    # err.error_code == "capacity_unavailable": nothing changed. Try again later.
```

From the command line: `shard operations wait <operation id>`.

**Starts wait for capacity for at most 15 minutes.** An open, resume or fork that no host can admit yet is
`capacity_pending`. If it is still pending 15 minutes after it was created, it fails with `capacity_unavailable` and
`retryable: true`: nothing was started, and a suspended workspace stays suspended with its state. The SDKs, the CLI
and the MCP server report this and do not retry by themselves.

A failed operation throws `OperationFailedError` (Python: raises) with `errorCode` / `error_code` and `retryable`.
Treat unknown error codes as generic errors: show the message and use `retryable`.

## Requests during transitions

| Request | While suspending | While suspended | While resuming |
| --- | --- | --- | --- |
| Tool call | Waits (the SDKs retry `workspace_busy`) | Wakes the workspace (SDKs, CLI, MCP); reads of files may be answered from its disk instead | Waits for the resume |
| Resume or open | `409 conflict` (`operation_in_progress`) | Resumes | Joins the resume |
| Suspend | Joins the suspend | `409 conflict` (`not_running`) | `409 conflict` (`operation_in_progress`) |

A refused call was never executed, so retrying it is safe. A command that is running when a suspend begins is frozen
with the VM and continues after the resume; its output stays readable by offset.

## Timings

Every open, wake and waited lifecycle call is timed. `formatTiming(workspace.lastTiming!)` (TypeScript),
`format_timing(ws.last_timing)` (Python) and `--timing` (CLI) print where the time went:

```text
open 34.18 s, succeeded (workspace 01a0e5a8-3edd-74ba-b489-d62b8925e342, operation 01a0e5a8-3ef0-7ecb-975e-dff2d5ca6e33)
  client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 590 ms → view 42 ms ∥ token 61 ms
  server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
  outside the server: 161 ms
```

This slow open spent its time waiting for a host with capacity (`capacity_pending`); starting the VM itself took
under a second. `outside the server` is network and polling time between you and the API. Watch the phases live with
`onProgress` (TypeScript) or `on_progress` (Python).
