Lifecycle
Open, suspend, resume, fork, reset and delete Shardflux workspaces: states, idle suspend and parking, waiting, and what survives each transition.
States and operations
A workspace's state (the API's observed_state) is one of:
| State | Meaning |
|---|---|
creating, starting |
The first open or a restart is booting it. |
running |
It runs, and tool calls are served. |
suspending |
Its memory, processes and disk are being checkpointed. Tool calls wait. |
suspended |
The checkpoint is stored durably and the VM is gone. No compute; the disk still counts as storage. |
resuming |
The checkpoint is being restored. |
forking |
A fork's copy is being created (the new workspace's state). |
failed |
A start failed. Opening the key again starts it again. |
deleting, deleted |
It is being or has been deleted. |
A running workspace that nobody uses is parked by its host. That is not a state:
it stays running. A file-first workspace is running from its first open until it is deleted.
Every change of state is an operation with a kind (open, suspend, resume, fork, snapshot, reset,
delete) and a state of its own: queued, capacity_pending, running, then succeeded, failed or canceled. A
workspace runs one lifecycle operation at a time; a second one is refused with 409 conflict
(details.reason: operation_in_progress), except that opening or resuming joins a start already in progress.
Requested or finished
Lifecycle calls return as soon as the change is requested, or, with wait, once it has finished:
| Call | Returns when | Returns |
|---|---|---|
await workspace.suspend() |
The suspend is requested (usually queued; the workspace still runs) |
The operation |
await workspace.suspend({ wait: true }) |
The suspend has finished; workspace.state is suspended |
The succeeded operation |
ws.suspend() (Python) |
The suspend is requested | The operation |
ws.suspend(wait=True) (Python) |
The suspend has finished; ws.state is suspended |
The succeeded operation |
shard ws suspend <key> |
The suspend is requested | Prints the operation |
shard ws suspend <key> --wait |
The suspend has finished | Prints the operation |
The same applies to resume, snapshot, delete, reset and close. open() waits by default. fork waits by
default in Python and takes { wait: true } in TypeScript.
Open, create or reconnect
open is the one call you need for most of the lifecycle. It creates the workspace for a new key, returns a running
one, resumes a suspended one, and joins a start already in progress. It never resets anything. See
What opening a key does.
Suspend a workspace
await workspace.suspend({ wait: true }); // memory and processes are checkpointed; compute stopsws.suspend(wait=True)A suspend pauses the VM, stores its memory and disk durably, and then releases the CPU and memory. The workspace is
suspended only after the checkpoint is stored. If storing it fails, the suspend fails and the workspace runs again
from its local copy; nothing is silently lost.
Session workspaces cannot be suspended (409, details.reason: session_lifetime).
Resume a workspace
You rarely need to resume explicitly:
- Opening the key resumes a suspended workspace.
- A tool call wakes it. An exec, file or other tool call on a suspended workspace resumes it (or joins the resume
already running), then runs. A call made during a suspend or resume waits for it to finish. The call never runs
twice. The SDKs bound the wait per call (default 120 seconds,
transitionTimeoutMsin TypeScript,transition_timeoutin Python); the CLI uses--wake-timeoutand the MCP serverSHARDFLUX_WAKE_TIMEOUT_MS. - The wake is one request (
@shardflux/sdk0.9.0+,shardflux0.5.0+ for Python, CLI 0.5.0+, MCP server 0.4.0+). The client sends one resume that the API holds until the workspace runs, and the answer carries a tool token for that client, so the refused call is retried at once. An explicitresume({ wait: true }),resume(wait=True)orshard ws resume --waitis the same single request. Earlier versions wait for the resume operation and then fetch a token. - Reads do not wake it. Reading, listing, stating and searching files of a suspended workspace are answered from
its disk while a host still holds it, without a resume (
@shardflux/sdk0.9.0+,shardflux0.5.0+, CLI 0.5.0+, MCP server 0.4.0+). See Search and edit files. - Following an exec's output never wakes a workspace, so a suspend you asked for is respected.
await workspace.resume({ wait: true }); // explicit
await workspace.wake(); // resume if needed; resolves once it runs
const cell = workspace.cell({ wake: null }); // opt out: calls fail with workspace_not_running insteadws.resume(wait=True)
ws.wake() # True if it resumed or waited, False if it was already running
cell = ws.cell(wake=None) # opt outA resume is admitted like a new start, so plan limits apply (see Pricing and limits).
Fork a workspace
A fork creates a new workspace under a new key from the current state of a running or suspended one: a copy of its disk and memory, restored into an independent VM with its own network identity. The source keeps running, or stays suspended.
const { workspace: copy } = await workspace.fork({ key: 'customer-42/experiment' }, { wait: true });copy = ws.fork("customer-42/experiment") # waits until the copy runsshard ws fork customer-42/main customer-42/experiment --wait- The new key must be unused, including by deleted workspaces (
409,details.reason: key_in_use). - The fork keeps the source's template version, secret bindings and text inputs. It gets the source's caps unless you
pass
caps. - A fork is persistent unless you pass
lifetime: 'session'. Forking is how you keep a session workspace's state. - A fork is a new start: plan limits apply to it as to an open.
Snapshot a workspace
workspace.snapshot({ label }) (TypeScript), ws.snapshot(label=...) (Python) records a checkpoint of a running or
suspended workspace without stopping it. While a running workspace is captured, reads work and writes wait. The API
has no call to list or restore snapshots yet; to branch from a state, fork.
Reset a workspace
reset wipes every change a workspace made and restarts it on its template version. It keeps the key, id, template
version, caps, secret bindings and volume attachments. Running processes end. It works only on layered workspaces
(workspace.diskLayout === 'layered'; others get 409, details.reason: legacy_disk_layout). The previous state is
kept as a recovery point for 7 days.
await workspace.reset({ wait: true });shard ws reset customer-42/main --yes --waitA suspended workspace stays suspended and starts blank on its next resume.
Delete a workspace
await workspace.delete({ wait: true });shard ws delete customer-42/main --yesTool access ends at once and the workspace's storage is cleaned up. Deletion cannot be undone, and a persistent workspace's key is never reused. The MCP server does not offer deletion to agents; use the SDK, the CLI or the console.
A session workspace ends with close() (shard ws close <key>), which deletes it; its key then opens a new
workspace.
Automatic suspend when idle
A persistent workspace suspends itself when it has been idle for its idle timeout. Nothing is killed: the suspend keeps memory and processes like any other. A running workspace uses RAM GiB-hours for every second it is awake, idle or not, so the shorter the idle tail after its last work, the less it costs you.
What keeps a workspace awake:
- a tool call (exec, files, terminal input, processes, git, browser);
- an attached exec output stream or terminal;
- a command started through exec, until it ends (for at most 1 hour, or its own timeout);
- a keepalive (
POST /v1/workspaces/{id}/keepaliveon the workspace's cell endpoint, for up to 12 hours).
A detached process, a dev server nobody is attached to, or CPU use alone does not keep it awake. Such processes continue after the resume.
Idle policy. The default policy is adaptive: each workspace learns its own timeout, from 10 seconds to 4 hours.
- It learns from the workspace's idle periods, every pause of 5 seconds or more between one activity and the next. Pauses inside its usual active hours and outside them are learned separately.
- Short pauses, such as an agent thinking between tool calls, are not worth a suspend and a resume, so the timeout outlasts them. When a workspace is woken soon after an automatic suspend, its next timeouts get longer (a suspend you asked for never has this effect).
- Until the workspace has 8 idle periods of its own, its history is blended with that of its template, else its organization, else all workspaces. With no history at all, the timeout is 5 minutes.
- The work signals above always win: the policy only decides how long an idle workspace waits.
The other policies are never and fixed:<seconds> (60 to 604800). The workspace view reports under idle the
policy, the current timeout (timeout_seconds), what it is based on (basis, for example learned from 37 idle periods (active-hours), template prior or default), when it would suspend (next_eligible_at) and a pending
suspend when idle request (suspend_request).
Set it with the HTTP API; the SDKs have no dedicated method yet, so use their request() passthrough:
await cloud.request('PUT', `/v1/workspaces/${workspace.id}/idle-policy`, { json: { idle_policy: 'fixed:600' } });sf.request("PUT", f"/v1/workspaces/{ws.id}/idle-policy", json={"idle_policy": "never"})null clears the workspace's own policy, so the template default (else adaptive) applies. Session workspaces have
no idle policy; they end after their idle timeout (10 minutes unless the template sets one).
Suspend when idle
Your code knows when an agent's turn ends. Instead of waiting for the idle timeout, ask for a suspend once the workspace has been idle for a short time. The request is stored with the workspace, so your process does not have to stay around for the suspend.
await workspace.suspendWhenIdle({ afterSeconds: 60 }); // 30 to 3600
workspace.suspendRequest; // { requested_at, after_seconds, not_before }, or null
await workspace.cancelSuspendWhenIdle(); // idempotentws.suspend_when_idle(after_seconds=60) # 30 to 3600
ws.suspend_request # SuspendRequest(requested_at, after_seconds, not_before), or None
ws.cancel_suspend_when_idle() # idempotentshard ws suspend customer-42/main --when-idle 1m # 30s to 1h
shard ws suspend customer-42/main --cancel-when-idle- The workspace is suspended once it has been idle for
after_seconds, counted from the later of its last work and the request.not_beforeis the earliest suspend. - A command still running, an attached exec output stream or terminal, or a keepalive postpones the suspend until
after_secondsafter it ends. The request stays. - The next tool call on the workspace (the next turn) or a resume cancels the request. Asking again replaces it.
- It applies under every idle policy,
neverincluded, and never delays a suspend the policy would do sooner. - If a suspend is already in progress, the call returns that operation and records nothing.
- A session workspace is refused (
409,details.reason: session_lifetime), and so is a workspace that is not running (not_running).
It is in @shardflux/sdk 0.10.0+ (cloud.workspaces.suspendWhenIdle(id, { afterSeconds }) by id), shardflux
0.6.0+ on PyPI (sf.workspaces.suspend_when_idle(workspace_id, after_seconds=60)), @shardflux/cli 0.5.1+ and
the MCP server 0.4.1+ (workspace_suspend with after_seconds). Over HTTP it is
POST /v1/workspaces/{id}/suspend-when-idle with {"after_seconds": 60}, and DELETE on the same path cancels it.
The workspace view shows a pending request as idle.suspend_request. Give your agent workspace
tools shows the call at the end of an agent loop.
Idle running workspaces are parked
Long before its idle timeout, a running workspace that nobody is using is parked by its host: after a few seconds
without activity its VM is paused and most of its memory is compressed, and after a longer pause the VM is saved to the
host's local disk. Parking is not a lifecycle state and needs nothing from you. The workspace stays running, keeps
its files, memory and processes, and the API answers as before. The next tool call wakes it first and then runs: about
a millisecond after a short pause, about a tenth of a second after a longer one.
- What keeps it resident: a tool call in progress, an attached exec output stream or terminal, a command started through exec until it ends, a keepalive, and processes that keep using CPU or the network.
- Background processes that wait (an idle dev server, a file watcher, a shell) are paused with the workspace and continue at the next wake. Timers inside the workspace fire late, never early, and the clock is right after the wake.
- Network traffic does not wake it. A process that only waits for data from the network is paused too, and a
remote peer that expects a prompt answer may time out. A keepalive (
POST /v1/workspaces/{id}/keepaliveon the cell endpoint) keeps the workspace resident, and wakes it if it is parked. - Billing does not change. A parked workspace is running: it uses RAM GiB-hours at its full memory allocation and counts toward running workspaces at once. CPU hours count the CPU time its processes use, which is next to none while it is parked. To stop RAM GiB-hours, suspend the workspace, or let the idle policy suspend it: parking neither delays nor replaces the automatic suspend.
- Reads may not wake it. Reading, listing, stating and searching files of a parked workspace can be answered from
its disk (
X-Served-From: disk) without waking it. - A wake can be refused for a moment. When the host has no room to restore the workspace right away, a tool call
is refused with
503 service_unavailable(details.reason: host_capacity) andRetry-After; when a restore fails, withwake_failed(the saved state is intact). Both are retryable, and nothing was executed. The SDKs retry reads, searches and calls with anIdempotency-Key; other calls surface the error withretryable: true.
Wake hint
A wake hint tells the host that a tool call is coming, so a parked workspace starts waking while your model is still writing the call:
void workspace.hint().catch(() => {}); // @shardflux/sdk 0.9.0+: returns at oncews.hint() # shardflux 0.5.0+: returns at once- A hint is cheap and never waits. For a suspended workspace it starts the resume in the background (TypeScript
result.wake, a promise; PythonWakeHint.wake, aFuture);hint({ wake: null })/hint(wake=False)only reports. - The TypeScript agent tools send a hint when each tool call starts, except for
read_file,list_filesandsearch_files, which a sleeping workspace answers from its disk.workspaceTools(ws, { hint: false })turns that off. The MCP server (0.4.0+) does the same. - Over HTTP,
POST /v1/workspaces/{id}/wake-hinton the cell endpoint answers202with theresidencythe host found:resident,frozen(paused),hibernated(saved to the host's disk) orrestoring. A suspended workspace answers409 workspace_not_running: resume it through the API. A hint that no tool call follows within 60 seconds is dropped. A hint is not tool activity and does not keep the workspace from its idle suspend.
What survives each transition
Files and disk, and memory and running processes, are listed separately.
| Transition | Files and disk | Memory and running processes | Also |
|---|---|---|---|
| Open the same key again | Kept | Kept | Nothing is reset. |
| Parked while idle, then woken by a tool call | Kept | Kept: processes are paused and continue | The workspace stays running and is billed as running. Timers fire late. |
| Suspend (by you, when idle, or at a compute allowance or spend cap), then resume | Kept: every file write acknowledged before the suspend began, and installed packages | Kept: every process with its PID, memory, open files, working directory and environment, including detached processes, dev servers and exec or terminal sessions; listening sockets and loopback connections | Your connections to exec output and terminals end; reattach by offset (the SDK does). Remote peers may close connections while the workspace sleeps. |
| Fork | Copied into the new workspace | Copied into the new workspace, as an independent VM with a new network identity | The source is unchanged. Shared volumes are attached, not copied: both see the same data. |
| Snapshot | Unchanged | Unchanged | Writes wait during the capture. |
| Reset (layered workspaces) | Wiped back to the template | Ended | Key, id, template version, caps, secrets and volumes are kept; the old state is a recovery point for 7 days. |
Session ends (close() or idle timeout) |
Deleted | Ended | The key opens a new, empty workspace. |
| Delete | Deleted | Ended | The key of a persistent workspace is never reused. |
A start fails with capacity_unavailable |
Unchanged | Unchanged | Nothing was started; a suspended workspace stays suspended. |
A resume always restores memory and disk from one checkpoint; it is never a fresh boot of the disk. Shared volumes are outside checkpoints: a suspend keeps the attachment, and the resume mounts the volume's current contents.
Wait for an operation to finish
wait takes options: { timeoutMs, signal, onProgress } in TypeScript (default 5 minutes), timeout= in seconds in
Python (default 300). When the time runs out, the SDK throws OperationTimeoutError, but the operation continues on
the server. Wait for it again:
import { OperationFailedError, OperationTimeoutError } from '@shardflux/sdk';
try {
await workspace.resume({ wait: { timeoutMs: 60_000 } });
} catch (err) {
if (err instanceof OperationTimeoutError) {
await cloud.workspaces.waitForOperation(err.operationId); // keep waiting
} else if (err instanceof OperationFailedError && err.retryable) {
// err.errorCode === 'capacity_unavailable': no host had room; nothing changed. Try again later.
} else {
throw err;
}
}from shardflux import OperationFailedError, OperationTimeoutError
try:
ws.resume(wait=True, timeout=60)
except OperationTimeoutError as err:
sf.workspaces.wait_for_operation(err.operation_id)
except OperationFailedError as err:
if not err.retryable:
raise
# err.error_code == "capacity_unavailable": nothing changed. Try again later.From the command line: shard operations wait <operation id>.
Starts wait for capacity for at most 15 minutes. An open, resume or fork that no host can admit yet is
capacity_pending. If it is still pending 15 minutes after it was created, it fails with capacity_unavailable and
retryable: true: nothing was started, and a suspended workspace stays suspended with its state. The SDKs, the CLI
and the MCP server report this and do not retry by themselves.
A failed operation throws OperationFailedError (Python: raises) with errorCode / error_code and retryable.
Treat unknown error codes as generic errors: show the message and use retryable.
Requests during transitions
| Request | While suspending | While suspended | While resuming |
|---|---|---|---|
| Tool call | Waits (the SDKs retry workspace_busy) |
Wakes the workspace (SDKs, CLI, MCP); reads of files may be answered from its disk instead | Waits for the resume |
| Resume or open | 409 conflict (operation_in_progress) |
Resumes | Joins the resume |
| Suspend | Joins the suspend | 409 conflict (not_running) |
409 conflict (operation_in_progress) |
A refused call was never executed, so retrying it is safe. A command that is running when a suspend begins is frozen with the VM and continues after the resume; its output stays readable by offset.
Timings
Every open, wake and waited lifecycle call is timed. formatTiming(workspace.lastTiming!) (TypeScript),
format_timing(ws.last_timing) (Python) and --timing (CLI) print where the time went:
open 34.18 s, succeeded (workspace 01a0e5a8-3edd-74ba-b489-d62b8925e342, operation 01a0e5a8-3ef0-7ecb-975e-dff2d5ca6e33)
client: request 20.01 s (held) → capacity_pending 13.52 s (no_ready_host) → running 590 ms → view 42 ms ∥ token 61 ms
server: queued 33.40 s, ran 620 ms, total 34.02 s; start warm, boot to ready 79 ms
outside the server: 161 msThis slow open spent its time waiting for a host with capacity (capacity_pending); starting the VM itself took
under a second. outside the server is network and polling time between you and the API. Watch the phases live with
onProgress (TypeScript) or on_progress (Python).