# Pricing and limits

> Shardflux plans and allowances, what happens at a limit, opt-in overage with a spend cap, workspace sizes, region, rate limits and file and command limits.

## Plans

Prices are in US dollars per month. Allowances reset with each billing period: your subscription's period on a paid
plan, the calendar month (UTC) on Free.

| | Free | Developer | Startup | Scale |
| --- | --- | --- | --- | --- |
| Price | $0 | $9 | $79 | $299 |
| CPU hours | As available | 100 | 400 | 1,000 |
| RAM GiB-hours | As available | 800 | 3,200 | 8,000 |
| Retained state | 10 GiB | 100 GiB | 1,000 GiB | 3,000 GiB |
| Workspaces stored (2 GiB each) | about 5 | about 50 | about 500 | about 1,500 |
| Outbound transfer | 10 GB | 100 GB | 1 TB | 5 TB |
| Running workspaces at once | 3 | 10 | 50 | 200 |
| Largest workspace (vCPU / RAM / disk) | 1 / 2 GiB / 2 GiB | 4 / 8 GiB / 20 GiB | 8 / 16 GiB / 50 GiB | 16 / 32 GiB / 100 GiB |
| Shared volumes (count / largest) | 1 / 1 GiB | 5 / 10 GiB | 20 / 100 GiB | 100 / 1,000 GiB |
| Overage spend cap per period (opt-in) | None | $1 to $9 | $1 to $79 | $1 to $299 |

**Free** has no guaranteed CPU or RAM hours: it runs on capacity as available. Its usage is measured and shown, but not
capped.

**Retained state** is sized for one workspace per customer, with suspended workspaces kept for as long as you want
them. The workspace counts assume 2 GiB each, which is generous for most work: on the default template with 2 GiB of
memory, a new workspace keeps under 0.1 GiB, one with a cloned Node.js repository and its dependencies about 0.2 GiB,
and one with a Python environment, 40 MB of data and a report about 0.4 GiB. A checkpoint also holds the memory the
workspace has used, so one that has filled its memory can keep close to its memory size on top of its disk.

The workspace counts are a storage estimate, not a limit: there is no cap on how many workspaces you create or keep,
and suspended workspaces do not count toward running workspaces at once.

**Enterprise**: contact [shardflux@heliosone.fi](mailto:shardflux@heliosone.fi).

There is no automatic overage. Reaching an allowance adds charges only after you turn on
[overage](#overage-opt-in) with a spend cap. Upgrade, change or cancel a plan under **Usage & billing** in the
[console](https://app.shardflux.dev), or from a terminal with the CLI (0.5.0+): `shard billing upgrade <plan>` prints
the Stripe Checkout URL where a person pays, and `shard billing portal` prints the link for plan changes and
cancellation ([Billing](https://docs.shardflux.dev/reference/cli.md#billing)).

## What each allowance measures

| Allowance | Measures |
| --- | --- |
| CPU hours | CPU time the workspace's processes actually used. |
| RAM GiB-hours | The workspace's memory allocation for as long as it runs: a 2 GiB workspace running for one hour uses 2 GiB-hours. A suspended workspace uses none. A running workspace that its host has [parked](https://docs.shardflux.dev/concepts/lifecycle.md#idle-running-workspaces-are-parked) while idle is still running and uses its full allocation. A [file-first workspace](https://docs.shardflux.dev/concepts/file-first.md) uses it only while an execution runs, at the execution VM's memory. |
| Retained state | Storage you keep: each workspace's own disk changes and checkpoints (running or suspended), a file-first workspace's current file tree, plus your organization's template storage. Platform templates never count. Measured as each day's peak, averaged over the billing period. |
| Outbound transfer | Bytes your workspaces send to the internet, in decimal GB. Inbound traffic is free and does not count, and neither does Shardflux's own traffic (checkpoints, template downloads, moves between hosts). |
| Running workspaces at once | Workspaces that are running or starting, across the whole organization. Suspended workspaces do not count, and file-first workspaces never count. |

Shared volume storage is measured and shown separately; no allowance includes it and it never blocks anything.

`shard usage`, `cloud.usage` (TypeScript) and **Usage** in the console show what you have used, with the time the
numbers were measured through.

## What happens at a limit

| Limit | What happens |
| --- | --- |
| Running workspaces at once | Opens, resumes and forks beyond the limit are refused with `403 quota_exceeded` (`details.limit: concurrent_workspaces`, `limit_value`, `current`). Suspend or delete a workspace, or upgrade. A workspace deleted a moment ago can still count for a short time. |
| CPU hours or RAM GiB-hours (paid plans), overage off | New opens, resumes and forks are refused with `402 allowance_exhausted` (`details`: `reason: allowance_used`, `allowance`, `used`, `included`, `resets_at`). Running workspaces are suspended, keeping their files, memory and processes. Nothing is deleted. |
| CPU hours or RAM GiB-hours (paid plans), overage on | Workspaces keep opening and running. Usage past the allowance is charged at the [overage rates](#overage-opt-in) until the charges reach your spend cap. At the cap, new opens, resumes and forks are refused with `402 allowance_exhausted` (`reason: spend_cap_reached`) and running workspaces are suspended the same way. Nothing is deleted. |
| Outbound transfer | At 100 %, outbound internet traffic of every workspace in the organization is blocked until you upgrade, add more, or the next period starts. Workspaces keep running; inbound connections, commands, terminals and file transfer keep working, and opens are not refused. |
| Retained state | Shown and notified when you are over it. It never blocks access or adds charges, and nothing is deleted. |
| Workspace size | Caps above the plan's maximum are lowered to it, not refused. |
| Template builds | 3 builds at once per organization; more are refused with `403 quota_exceeded` (`concurrent_template_builds`). A version larger than the plan's disk maximum fails with `resource_limit`. |
| Shared volumes | Creating one beyond the count or size is refused with `403 quota_exceeded` (`shared_volumes_max`, `shared_volume_gib_max`). At most 8 volumes can be attached to one workspace. |

Every owner and billing member gets an email at 80 % and 100 % of an allowance. With overage on, they also get one at
50 %, 80 % and 100 % of the spend cap. Compute limits are enforced with short budget leases, so usage can run slightly
past an allowance or the spend cap before running workspaces are suspended; the organization's spend view reports
that bound.

**Billing states.** If a payment fails, you have 7 days of grace. Nothing changes during it, except that
[overage](#overage-opt-in) is paused until the invoice is paid. After that, new opens,
resumes and forks are refused with `402 entitlement_required` (`details.reason: payment_past_due` or `unpaid`); running
workspaces and your data are untouched. A cancellation takes effect at the end of the period, after which Free limits
apply to new starts.

**Data is never deleted by billing or inactivity.** Workspaces are deleted only when you delete them, with one
exception: a [session workspace](https://docs.shardflux.dev/concepts/workspaces.md#persistent-and-session-workspaces) is deleted when its session
ends.

## Overage (opt-in)

Overage keeps your workspaces opening and running past the CPU-hours and RAM GiB-hours allowances, up to a spend cap
you set. It is off by default. Free has no overage.

**Turning it on.** An owner or billing member turns overage on under **Usage & billing** in the
[console](https://app.shardflux.dev) and sets a spend cap in dollars per billing period. The cap is at least $1 and at
most the plan price: $9 on Developer, $79 on Startup, $299 on Scale. A period's bill is therefore at most twice the plan
price. The cap never exceeds the current plan's price, so moving to a cheaper plan lowers it.

**Rates.** They are the same on every paid plan.

| Past the allowance | Rate |
| --- | --- |
| RAM GiB-hours | $0.04 per GiB-hour. A 2 GiB workspace awake for an hour costs $0.08. |
| CPU hours | $0.12 per CPU-hour past the plan's CPU hours. |

Storage and outbound transfer have no overage. Retained state and shared volumes are never charged, and outbound
traffic is still blocked at 100 % of the transfer allowance.

**Billing.** Charges come from the same usage measurements as the allowances, counted per UTC hour. They go on your
next monthly invoice and are charged to the card that pays the plan. There are no prepaid credits. Upgrading during a
period does not refund overage already charged.

**At the cap.** New opens, resumes and forks are refused with `402 allowance_exhausted`
(`details.reason: spend_cap_reached`; see [Errors](https://docs.shardflux.dev/reference/errors.md#402-allowance_exhausted)). Running workspaces are
suspended, keeping their files, memory and processes. Nothing is deleted. Work can start again when the next billing
period begins, when you raise the cap (up to the plan price), or when you upgrade. Usage can run slightly past the cap
before workspaces are suspended, but you are never charged more than the cap.

**Past-due payments.** While a plan payment is past due, overage is paused and the allowances apply as if it were off.
Past an allowance, starts are refused with `402 allowance_exhausted` (`details.reason: overage_paused`) and running
workspaces are suspended. Overage resumes when the invoice is paid.

**Emails.** Every owner and billing member gets an email at 50 %, 80 % and 100 % of the cap, and when someone turns
overage on or off or changes the cap.

## Workspace sizes

A workspace's CPU, memory and disk are the smallest of three maximums: the plan's (above), the template's, and the caps
you pass to `open` (`caps: { cpu_millis, memory_mib, disk_gib }`; 1,000 `cpu_millis` is one vCPU). The allocation is
fixed while the workspace runs.

- `python-node-browser` allows up to 4 vCPU, 8 GiB of memory and 20 GiB of disk, and starts workspaces with 2 GiB of
  memory unless you set `caps.memory_mib`.
- Disk is the space for the workspace's own changes; the template's files do not count against it.

See [Workspace size](https://docs.shardflux.dev/concepts/workspaces.md#workspace-size).

## Region

Workspaces and the API run in AWS `us-west-2` (Oregon, United States). The API is `https://api.shardflux.dev` and the
console `https://app.shardflux.dev`.

## Starts and waits

| | Limit |
| --- | --- |
| A start waiting for capacity (`capacity_pending`) | At most 15 minutes, then `capacity_unavailable` (retryable; nothing started) |
| SDK wait for an operation | 5 minutes by default (`timeoutMs`, `timeout`); the operation continues after the SDK stops waiting |
| Wake of a suspended workspace by a tool call | 120 seconds by default |
| Idle timeout of a persistent workspace | Learned per workspace, 10 seconds to 4 hours; 5 minutes while there is no history to learn from. Or `fixed:<seconds>` (60 to 604,800) or `never`. See [Automatic suspend when idle](https://docs.shardflux.dev/concepts/lifecycle.md#automatic-suspend-when-idle) |
| Suspend when idle | `after_seconds` 30 to 3,600 ([Suspend when idle](https://docs.shardflux.dev/concepts/lifecycle.md#suspend-when-idle)) |
| Session idle timeout | 10 minutes unless the template sets one (60 to 86,400 seconds) |

## API rate limits

| Limit | Value |
| --- | --- |
| Tool tokens (including an open that returns one) | 3,600 per hour per API key |
| Template package searches | 120 per minute per organization |

Over a limit, the API answers `429 rate_limited` with a `Retry-After` header. The SDKs retry `429` for safe and
idempotent requests (`maxRetries` / `max_retries`, default 2), then raise the error with `retryAfterSeconds`
(TypeScript) or `retry_after` (Python).

## File and command limits

| | Limit |
| --- | --- |
| One file write | 256 MiB (`413 payload_too_large` above) |
| One file read | 256 MiB per request; the SDKs read larger files in several requests |
| File paths | Absolute, at most 4,096 bytes |
| Directory listing | Up to 10,000 entries per call (default 1,000) |
| Command arguments | Up to 4,096 arguments of up to 128 KiB each |
| Command environment | Up to 512 variables |
| Command stdin | 1 MiB |
| Command timeout | Up to 7 days |
| Command output kept by `exec.run` / `ws.exec` | 1 MiB of stdout and 1 MiB of stderr by default (`maxOutputBytes` / `max_output_bytes`); stream everything with `onOutput` / `on_output` |
| Output returned to a model by the agent tools | 64 KiB per call |
| Reads of a suspended workspace from its disk | The same limits as other reads; not billed as compute |
| Secrets bound to one workspace | 50 |
| Browser screenshot / page content | 32 MiB / 5 MiB |
| Workspace keys | 1 to 200 characters; keys starting with `sf:` are reserved |

### Search, patches and revisions

| | Limit |
| --- | --- |
| Search pattern | 1 to 1,000 characters, literal or RE2 |
| Search results | 200 matches by default, at most 5,000; a search also stops after 10 seconds or 4 MiB of results (`truncated`, `stop_reason`) |
| Files searched | Up to 1 MiB each by default, at most 64 MiB (`max_file_bytes`); each line in its first 1 MiB; binary files and symbolic links are skipped |
| Search globs | Up to 32 `include` and 32 `exclude` globs of up to 256 characters |
| Search context | 0 to 5 lines before and after each match; the matched line is returned up to 1,000 bytes |
| One patch request | 7 MiB (`413 payload_too_large` above); write larger files with a normal write |
| A patched file | 64 MiB |
| Edits in one patch | 1 to 100 |
| File revision from `stat` | Regular files up to 256 MiB |
| File revision on a read (`X-File-Revision`) | Regular files up to 16 MiB |
| Wake hint | Held for 60 seconds when no tool call follows |

See [Search and edit files](https://docs.shardflux.dev/guides/files.md).

### File-first workspaces

| | Limit |
| --- | --- |
| Files | Only under `/home/user`; names must be valid UTF-8 |
| Executions at once | One per workspace; others and file changes get `409 workspace_busy` (`execution_in_progress`) until it ends |
| Execution output | stdout and stderr each up to 1 MiB by default, at most 16 MiB (`outputLimitBytes`, `output_limit_bytes`, `--output-limit`) |
| Changed paths in a result | Up to 10,000 (`changed_truncated`); the agent tools return up to 200 |
| Execution ids | 8 to 128 characters of `A-Z a-z 0-9 . _ : -`, starting with a letter or a digit |
| Execution results | Kept for 7 days |
| Command stdin, timeout, arguments and environment | As for other commands (above) |
| Lifetime | Persistent only |

See [File-first workspaces](https://docs.shardflux.dev/concepts/file-first.md).
