Give your agent workspace tools
Give your agent a Shardflux workspace: workspace tools in TypeScript and Python for Anthropic, OpenAI and other frameworks, scheduled runs, tool-call capture.
Your agent loop, the workspace's tools
Your application keeps the agent loop and the model calls. The workspace is the computer the agent's tools act on.
workspaceTools(workspace) from @shardflux/sdk returns the tools: each has a name, a description, a JSON Schema
for its parameters and an execute function that runs the call in the workspace.
import { executeToolCall, toAnthropicTools, toOpenAITools, workspaceTools } from '@shardflux/sdk';
const tools = workspaceTools(workspace);
const anthropicTools = toAnthropicTools(tools); // Anthropic Messages API
const chatTools = toOpenAITools(tools); // OpenAI Chat Completions
const responsesTools = toOpenAITools(tools, { api: 'responses' }); // OpenAI Responses API
// For each tool call the model makes:
const output = await executeToolCall(tools, { name: call.name, input: call.input });executeToolCall finds the tool by name, accepts the arguments as an object (input, Anthropic) or a JSON string
(arguments, OpenAI), validates them against the tool's schema, and runs the tool. An Anthropic tool_use block or an
OpenAI Responses function_call item can be passed as it is. (0.8.0+) The exported definitions are typed as the
providers' own tool types (Anthropic.Tool[], OpenAI.Chat.ChatCompletionTool[] and, with { api: 'responses' },
OpenAI.Responses.FunctionTool[]), so they type-check under TypeScript's strict checks without casts. The workspace keeps its files,
installed packages and processes between calls and between conversations. If the workspace is suspended, the first
tool call resumes it.
In Python (shardflux 0.4.0+), workspace_tools(ws) returns the same tools with the same names and schemas,
except search_files and edit_file (TypeScript only so far):
from shardflux import execute_tool_call, to_anthropic_tools, to_openai_tools, workspace_tools
tools = workspace_tools(ws)
anthropic_tools = to_anthropic_tools(tools) # Anthropic Messages API
chat_tools = to_openai_tools(tools) # OpenAI Chat Completions
responses_tools = to_openai_tools(tools, api="responses") # OpenAI Responses API
# For each tool call the model makes: a tool_use block, an OpenAI tool call or function_call item, or a dict.
output = execute_tool_call(tools, block)The Python tools are synchronous; in async code, run them with asyncio.to_thread. A coding agent such as Claude
Code gets the same tools from the MCP server.
A complete loop with Anthropic
npm install @shardflux/sdk @anthropic-ai/sdk
export ANTHROPIC_API_KEY=...import Anthropic from '@anthropic-ai/sdk';
import { Shardflux, executeToolCall, toAnthropicTools, workspaceTools } from '@shardflux/sdk';
const cloud = new Shardflux({ apiKey: process.env.SHARDFLUX_API_KEY! });
const anthropic = new Anthropic(); // reads ANTHROPIC_API_KEY
const workspace = await cloud.workspaces.open({ key: 'agent-demo/main', template: 'python-node-browser' });
const tools = workspaceTools(workspace, { tools: ['exec', 'files'] });
const messages: Anthropic.MessageParam[] = [
{ role: 'user', content: 'Write /home/user/fizzbuzz.py, run it for 1 to 15, and tell me what it printed.' },
];
for (;;) {
const response = await anthropic.messages.create({
model: 'claude-opus-5',
max_tokens: 16000,
tools: toAnthropicTools(tools),
messages,
});
messages.push({ role: 'assistant', content: response.content });
if (response.stop_reason !== 'tool_use') {
for (const block of response.content) if (block.type === 'text') console.log(block.text);
await workspace.suspendWhenIdle({ afterSeconds: 60 }); // the turn is over: suspend after a minute idle
break;
}
// Run every tool call of this turn, and send all results back in one message.
const results: Anthropic.ToolResultBlockParam[] = [];
for (const block of response.content) {
if (block.type !== 'tool_use') continue;
try {
const output = await executeToolCall(tools, block);
results.push({ type: 'tool_result', tool_use_id: block.id, content: JSON.stringify(output) });
} catch (err) {
results.push({ type: 'tool_result', tool_use_id: block.id, content: String(err), is_error: true });
}
}
messages.push({ role: 'user', content: results });
}Run it with node agent.ts (in a package with "type": "module"). Errors go back to the model as is_error results,
so it can correct a bad argument or a failing command. When the model answers without a tool call, the turn is over,
and the loop asks for a suspend once the workspace has been idle for a minute (0.10.0+); see
Suspend when the turn ends.
A complete loop with OpenAI
npm install @shardflux/sdk openai
export OPENAI_API_KEY=... OPENAI_MODEL=... # the OpenAI model you useWith the Responses API, a function_call item can be passed to executeToolCall as it is. This loop continues the
conversation with previous_response_id, so each turn sends only the tool results:
import OpenAI from 'openai';
import { Shardflux, executeToolCall, toOpenAITools, workspaceTools } from '@shardflux/sdk';
const cloud = new Shardflux({ apiKey: process.env.SHARDFLUX_API_KEY! });
const openai = new OpenAI(); // reads OPENAI_API_KEY
const workspace = await cloud.workspaces.open({ key: 'agent-demo/main', template: 'python-node-browser' });
const tools = workspaceTools(workspace, { tools: ['exec', 'files'] });
let input: OpenAI.Responses.ResponseInput = [
{ role: 'user', content: 'Write /home/user/fizzbuzz.py, run it for 1 to 15, and tell me what it printed.' },
];
let previousResponseId: string | undefined;
for (;;) {
const response = await openai.responses.create({
model: process.env.OPENAI_MODEL!,
tools: toOpenAITools(tools, { api: 'responses' }),
input,
previous_response_id: previousResponseId,
});
const calls = response.output.filter((item) => item.type === 'function_call');
if (calls.length === 0) {
console.log(response.output_text);
await workspace.suspendWhenIdle({ afterSeconds: 60 }); // the turn is over: suspend after a minute idle
break;
}
previousResponseId = response.id;
input = [];
for (const call of calls) {
const output = await executeToolCall(tools, call).catch((err: unknown) => ({ error: String(err) }));
input.push({ type: 'function_call_output', call_id: call.call_id, output: JSON.stringify(output) });
}
}For Chat Completions, pass toOpenAITools(tools) as tools and dispatch each function call of message.tool_calls
(call.type === 'function') with
executeToolCall(tools, { id: call.id, name: call.function.name, arguments: call.function.arguments }). A message
without tool_calls ends the turn: call workspace.suspendWhenIdle({ afterSeconds: 60 }) there.
A complete loop in Python
(shardflux 0.4.0+; suspend_when_idle 0.6.0+) The same Anthropic loop with the Python SDK:
pip install shardflux anthropic
export SHARDFLUX_API_KEY=... ANTHROPIC_API_KEY=...import json
import anthropic
from shardflux import Shardflux, execute_tool_call, to_anthropic_tools, workspace_tools
sf = Shardflux() # reads SHARDFLUX_API_KEY
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
ws = sf.open(key="agent-demo/main", template="python-node-browser")
tools = workspace_tools(ws, tools=["exec", "files"])
messages = [
{"role": "user", "content": "Write /home/user/fizzbuzz.py, run it for 1 to 15, and tell me what it printed."}
]
while True:
response = client.messages.create(
model="claude-opus-5", max_tokens=16000, tools=to_anthropic_tools(tools), messages=messages
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
print("".join(block.text for block in response.content if block.type == "text"))
ws.suspend_when_idle(after_seconds=60) # the turn is over: suspend after a minute idle
break
# Run every tool call of this turn, and send all results back in one message.
results = []
for block in response.content:
if block.type != "tool_use":
continue
try:
output = execute_tool_call(tools, block)
results.append({"type": "tool_result", "tool_use_id": block.id, "content": json.dumps(output)})
except Exception as err: # bad arguments or a refused call: tell the model, so it can correct itself
results.append({"type": "tool_result", "tool_use_id": block.id, "content": str(err), "is_error": True})
messages.append({"role": "user", "content": results})With the OpenAI Responses API, pass to_openai_tools(tools, api="responses") as tools, pass each function_call
item of response.output to execute_tool_call as it is, and send the result back as
{"type": "function_call_output", "call_id": item.call_id, "output": json.dumps(output)}. When response.output has
no function_call item, the turn is over: call ws.suspend_when_idle(after_seconds=60). A Chat Completions tool call
(message.tool_calls[i]) can also be passed as it is.
Suspend when the turn ends
A running workspace uses RAM GiB-hours and a running slot for every second it is awake, including the idle time before
its idle timeout suspends it. Your loop knows when a turn ends, so
the loops above end each turn with workspace.suspendWhenIdle({ afterSeconds }) (@shardflux/sdk 0.10.0+) or
ws.suspend_when_idle(after_seconds=...) (shardflux 0.6.0+): once the workspace has been idle for
afterSeconds (30 to 3600), it is suspended.
- If the user answers within that time, the next turn's first tool call cancels the request and finds the workspace running. Otherwise the workspace suspends, and the next tool call wakes it.
- A command the agent left running (counted for at most 1 hour, or its own timeout), an attached output stream or a
keepalive postpones the suspend until
afterSecondsafter it ends. - The idle time counts from the later of the workspace's last work and the request. Calling it again replaces the
request, and
cancelSuspendWhenIdle()(cancel_suspend_when_idle()) cancels it. - It never delays a suspend the idle policy would do sooner, and it applies under every policy,
neverincluded. - Choose
afterSecondsby how quickly your users usually reply: long enough to cover a quick answer, short enough that an abandoned conversation stops costing.
A job that nobody is waiting on does not need the grace period: suspend it at once (see Run the agent on a schedule or an event). All the rules are in Suspend when idle.
Other agent frameworks
The tools are plain objects, so any framework that takes a name, a description, a JSON Schema and a function can use them:
import type { WorkspaceTool } from '@shardflux/sdk';
const tools: WorkspaceTool[] = workspaceTools(workspace);
for (const t of tools) console.log(t.name, t.permission, JSON.stringify(t.parameters));
const result = await tools.find((t) => t.name === 'exec')!.execute({ command: 'python3 --version' });execute validates its arguments against the same schema the model saw and throws ToolArgumentError (with
issues) when they do not match. In Python, a WorkspaceTool has the same fields; call tool.execute({"command": "python3 --version"}). Its ToolArgumentError is also a ValueError.
The tools
Which tools you get depends on the API key's tool permissions and the tools option.
| Tool | Permission | Parameters | Returns |
|---|---|---|---|
exec |
exec |
command (run with bash -lc), cwd (absolute; a relative one is refused with invalid_cwd), timeout_ms (1,000 to 3,600,000; default 600,000), stdin |
exit_code, term_signal, timed_out, stdout, stderr, truncated, session_id; error (code, message, reason) when the command could not start (@shardflux/sdk 0.10.0+, Python 0.6.0+) |
read_file |
files |
path, offset, length |
path, content (UTF-8), truncated |
write_file |
files |
path, content, append, create_parents (default true) |
path, bytes_written, sha256, durable |
list_files |
files |
path, limit (up to 10,000; default 500) |
entries (name, path, type, size, modified_at), truncated |
search_files (TypeScript 0.9.0+, Python 0.6.0+) |
files |
path (a directory or one file), pattern (literal text, or RE2 with regex), regex, case_insensitive, include and exclude (globs; exclude defaults to .git and node_modules), max_matches (up to 5,000; default 200), context_lines (up to 5) |
matches (path, line, column, text, before, after), truncated, stop_reason, omitted_matches, files_scanned |
edit_file (TypeScript 0.9.0+, Python 0.6.0+) |
files |
path, edits (1 to 100 of old_text, new_text, replace_all), expected_revision |
path, revision, previous_revision, replacements, bytes_written |
list_processes |
process |
none | processes (pid, ppid, comm, cmdline, state, rss_bytes) |
signal_process |
process |
pid, signal (such as SIGTERM) |
signalled |
terminal_open |
pty |
command (default: a login shell), rows, cols |
session_id, state, next_offset |
terminal_send |
pty |
session_id, input (include \n to press Enter) |
state, next_offset |
terminal_read |
pty |
session_id, offset, wait_ms (up to 60,000) |
output, next_offset, exited, exit_code; Python also truncated |
terminal_close |
pty |
session_id |
state |
git_clone |
git |
url (HTTPS), path, branch, depth |
exit_code, stdout, stderr |
git_status |
git |
path |
The repository's status |
git_commit |
git |
path, message (stages all changes) |
exit_code, commit, stdout, stderr |
browser_screenshot |
browser |
url, width, height |
mime_type (image/png), bytes, data_base64 |
browser_content |
browser |
url, format (text or html) |
url, format, content, truncated |
Command output and file content returned to the model are cut at 64 KiB per call (truncated: true); maxOutputBytes
(max_output_bytes in Python) changes that. search_files returns whole matches within that limit and counts the
rest in omitted_matches. edit_file applies its edits only to the file's revision: the expected_revision the model
gives (the revision of its previous edit of that file), else the revision the tool reads just before editing, so a
change made in between fails with revision_mismatch instead of being overwritten. The Python SDK does not have these
two tools yet; see Search and edit files. In Python, the cut never splits a UTF-8 character, and terminal_read's
next_offset counts exactly the bytes it returned, so output past the limit comes with the next read.
search_files returns whole matches up to that limit and counts the rest in omitted_matches.
edit_file replaces exact text: each old_text must occur exactly once in the file unless replace_all is true, and
the edits apply in order, all of them or none. When the model gives no expected_revision, the tool reads the file's
revision first and pins the edit to it, so a change made in between fails the edit (409 conflict, reason: revision_mismatch) instead of being overwritten. The result's revision can be passed as expected_revision to the
next edit_file of the same file.
(TypeScript 0.9.0+, Python 0.6.0+) Each tool call first tells the workspace that a call is coming (a wake hint,
sent without waiting for it), so a workspace its host parked while idle is being restored while the call is prepared.
read_file, list_files and search_files send no hint, so reading a suspended workspace does not resume it through
the hint. Pass hint: false (hint=False in Python) to turn it off, for example when you send workspace.hint()
yourself as soon as the model starts a tool call.
Options
const tools = workspaceTools(workspace, {
tools: ['exec', 'files', 'git'], // default: the tools of the workspace's last token, else all six
agentLabel: 'coder', // attribution: one agent session per label
prefix: 'workspace_', // tool names become workspace_exec, workspace_read_file, ...
maxOutputBytes: 65_536, // bytes of output or file content returned to the model (default 64 KiB)
defaultCwd: '/home/user/project', // exec's working directory when the model gives none (default: the user's home)
transitionTimeoutMs: 120_000, // longest wait per call for a suspended workspace to wake (default 120 s)
hint: true, // (0.9.0+) send a wake hint when a call starts (default true)
});Pass wake: null to make calls on a suspended workspace fail with workspace_not_running instead of resuming it.
(0.9.0+) Each tool call first sends a wake hint without waiting for it, so a
workspace its host has parked while idle starts waking while the call runs (read_file, list_files and
search_files send none: a sleeping workspace answers them from its disk). hint: false turns that off, for example
when you send workspace.hint() yourself as soon as the model starts writing a tool call. On a file-first
workspace the tools are exec and the files tools only, and exec runs each command as an
execution: its result adds execution_id, state, tree_revision and changed. mode builds the definitions
without reading the workspace, and onExecution(id) is called with each execution id before it is sent.
In Python the options are keyword arguments with the same meaning: workspace_tools(ws, tools=[...], agent_label=..., prefix=..., max_output_bytes=..., default_cwd=..., transition_timeout=120, wake=None, hint=True) (hint from 0.6.0).
Run the agent on a schedule or an event
Scheduled and event-driven work starts where your agent loop runs: in your application. Have the job runner or webhook handler you already use (cron, Vercel Cron, Inngest, Trigger.dev, Celery beat, GitHub Actions) call your agent, and open the customer's workspace by key as usual. A suspended workspace resumes on the open, so nothing has to stay running between runs. Shardflux does not run schedules itself, and a cron job inside a workspace does not run while the workspace is suspended (see what keeps a workspace awake).
// Called by your job runner, once per customer.
export async function weeklyReport(customerId: string) {
const workspace = await cloud.workspaces.open({ key: `customer/${customerId}`, template: 'python-node-browser' });
try {
await runAgent(workspace, 'Update the weekly report with the new files in /home/user/data.'); // your loop
} finally {
await workspace.suspend({ wait: true }); // frees its running slot now, not at its idle timeout
}
}def weekly_report(customer_id: str) -> None:
ws = sf.open(key=f"customer/{customer_id}", template="python-node-browser")
try:
run_agent(ws, "Update the weekly report with the new files in /home/user/data.") # your loop
finally:
ws.suspend(wait=True)When one schedule covers many customers:
- Keep concurrency below your plan's running limit. Running workspaces count across your organization, including
the ones your users are working in. Past the limit, opens are refused with
403 quota_exceeded(details.limit: concurrent_workspaces), which the SDKs do not retry. Run the jobs through a queue with a concurrency limit, for example 40 on Startup (limit 50), and let the queue retry a refused job later. - Suspend when the job is done, with
suspend(). Nobody is waiting for a reply, so free the running slot at once.suspendWhenIdleis for interactive turns, where the user may answer within a minute. With neither, each workspace keeps running until its learned idle timeout (5 minutes while there is no history to learn from), using RAM GiB-hours for no work: 200 workspaces of 2 GiB that each stay awake 5 minutes after a daily job use about 1,000 GiB-hours a month, almost a third of Startup's 3,200. If one of your users makes a tool call during the suspend, the call waits and then wakes the workspace again. - Or keep only files. When a job needs nothing but its files between runs, a file-first workspace never holds a running slot, and compute is metered only while a command runs.
- Make jobs safe to retry. A start that finds no host with room fails after 15 minutes with
capacity_unavailable(retryable: true) and changes nothing. Opening a key never resets its workspace, so running a job again is safe as long as your own steps are.
See Pricing and limits for each plan's running limit and allowances.
Save your harness's tool calls into the workspace
Your own tools (web search, SQL, HTTP APIs, other MCP servers) run in your application, so their results reach the
model but not the workspace. Tool-call capture (@shardflux/sdk 0.7.0+, shardflux 0.3.0+) saves each call's input
and full output as files in the workspace, where the agent can work on them with jq or Python, and where snapshots and
forks keep them.
const capture = workspace.captureToolCalls();
const tools = capture.tools(workspaceTools(workspace)); // Shardflux tools are recorded too
// Inside your loop, for each tool_use block of the model's reply:
const output = myTools[block.name]
? await capture.run(block, () => myTools[block.name]!(block.input)) // your tool: recorded, same return value
: await executeToolCall(tools, block);
// Before your process or serverless function ends:
await capture.flush();In Python (0.4.0+): tools = capture.tools(workspace_tools(ws)), then execute_tool_call(tools, block) records
the call with block.id.
Capture is invisible to your harness: a wrapped tool returns the same value and throws the same error, and nothing
capture does throws into your code (failures go to onError and capture.stats). Writes run in the background.
In the workspace, under /home/user/tool-calls/<run>/:
index.jsonl one JSON line per call: seq, call_id, tool, status, input, output_path, ...
000007-web_search.json a call's output (.json, .txt, .html, .png, .pdf, ... from its content)
000010-github.search/ an MCP result or content blocks: part-1.txt, part-2.png, result.jsoncapture.promptHint() returns a paragraph telling the agent where its tool calls are; add it to your system prompt if
you want. The agent can then run, for example, jq -cR 'fromjson? // empty' /home/user/tool-calls/*/index.jsonl.
| Harness | TypeScript | Python |
|---|---|---|
| Hand-rolled loop | capture.run(call, fn), capture.wrap(name, fn), capture.record(...) |
with capture.call(name, input, call_id=...) as c: c.output = ..., capture.record(...) |
| Decorated tools | captureTool(name, fn) with capture.activate(fn) |
@capture_tool under the framework's decorator, with capture.activate() |
| Anthropic tool runner | capture.anthropic.tools([...]) |
@capture_tool under @beta_tool |
| OpenAI Agents | capture.openaiAgents.attach(runner) |
CaptureRunHooks(capture) (shardflux[openai-agents]) |
| Claude Agent SDK | capture.claude.hooks() |
capture_hooks(capture) (shardflux[claude-agent-sdk]) |
| Vercel AI SDK 7 | capture.aiSdk.tools(tools) |
- |
| Mastra | capture.mastra.tools(...), capture.mastra.hooks() |
- |
| LangChain / LangGraph | capture.langchain.handler() |
CaptureCallbackHandler(capture) (shardflux[langchain]) |
| Pydantic AI | - | capture_capability(capture) (shardflux[pydantic-ai]) |
| CrewAI | - | register_hooks(capture) (shardflux[crewai]) |
| MCP client | capture.mcp.instrument(client, { server }) |
instrument(session, capture, server=...) |
A Python hand-rolled loop:
capture = ws.capture_tool_calls()
for block in response.content:
if block.type == "tool_use":
with capture.call(block.name, block.input, call_id=block.id) as call:
call.output = my_tools[block.name](**block.input)
capture.flush()- Redaction. Nothing is redacted by default. Pass
transformto change or drop a call before it is written; if it throws, the call is dropped, never written unredacted. - Read your own writes. Exec and file calls through the same client, and suspend, snapshot and fork, first wait (up to 30 seconds) for captured writes recorded before them.
- Limits. An output over 32 MiB is cut (text and JSON) or not stored (binary). Pending writes are bounded, and a tool call never waits for capture. Around 100 calls per second per workspace reaches the pending limit.
- Serverless. Flush before the function returns:
waitUntil(capture.flush())on Vercel,await capture.flush()on AWS Lambda.
Every option is in the TypeScript and Python references.