← Sim Desk / API
Tokens

Drive Sim Desk from your own code

A reading of the model, not of the real system. The model reads the replication analysis your browser (or your own code) computed; it never recomputes a number. A queue simulation says what this model does under these assumptions (exponential arrivals and service, first-come first-served, one pool of identical servers), not what the real clinic, line or call centre will do.

Everything the web page does is available over HTTP. The page runs, in your browser, an exact port of the @k-dense-ai/simpy skill's bundled queue model and replication runner (SimPy 4.1.2 semantics): a finite-horizon M/M/c/K-style queue with servers, queue_capacity, exponential mean_interarrival and mean_service, a horizon, analysis_mode terminating or steady_state with a warm_up, a base_seed, replications and a confidence level. From the replications it derives Student-t intervals, the precision against a target, the exact M/M/c/K steady-state benchmark, a windowed view of the queue over time and, when you add a scenario B, paired B minus A differences with common random numbers. You send those facts and get back a verdict (sound, caveated, unreliable) and either a reading of every metric and the comparison, or a SimPy script that reproduces every number and applies the fixes. The model behind the run is gpt-terra. The natural loop: simulate, read, script, add replications or a longer warm-up, re-check.

Two lanes: the task field

taskwhat you getextra input
interpretA reading of each metric (M1.. in the order of facts.metrics: offered load per server, the interval metrics, unfinished customers, the worst relative half-width, the M/M/c/K references and, with a scenario B, the B minus A rows), what the study design estimates, the comparison call, your claims judged against the intervals, and what the study cannot show.none
scriptThe fixes (more replications from precision.rows, a longer horizon, a longer warm-up in a steady-state study, raising max_entities or max_events, running scenario B) and one complete Python script: the configuration constants, an EXPECTED dict of every browser value, the skill's BLAKE2b seed rule, the SimPy model and replication loop, a math.isclose check of every key, then the fixes, all inside main().decision: the text of an earlier interpret run (optional)

Both lanes return the same envelope: lane, verdict, headline, tldr, the lane body, next_steps and prescan_responses. Worked requests: interpret, script. The reply shape: output contract.

Input fields

Every field is a string.

fieldrequiredmeaning
taskyesinterpret or script.
factsyesA JSON-encoded string holding the browser's replication analysis - see below. The page builds it with DeskKit.buildInput; an API caller builds it too.
titlenoA label for the study, up to 160 characters.
contextnoYour notes: what the servers and customers stand for, the time unit, the decision the study informs, what you believe. Up to 3,000 characters.
decisionscript onlyPlain text of an earlier interpret run (the page builds it with Recon.decisionText: "Verdict: ...", the headline, one line per metric reading, the comparison call, then the next steps). Up to 5,000 characters.
questionnoAnswered in tldr as a bullet starting "Answer:". Up to 1,200 characters.
retry_notenoOnly on a retry after a malformed reply.

The facts string

facts is a JSON string, not an object: the browser runs the replications, serialises the analysis with JSON.stringify and sends that text. It holds: settings (time_unit, target_relative_half_width, simpy_version 4.1.2, schema_version, seed_rule); config (model with analysis_mode, horizon, warm_up, servers, queue_capacity, mean_interarrival, mean_service, base_seed as a string, max_entities, max_events; then replications and confidence); offered (arrival_rate, service_rate, offered_load, rho_per_server, system_capacity); totals (arrivals, rejected, admitted, completed, observed_completed, unfinished, replications_with_unfinished, entity_limit_hits, no_completions, max_events_used, event_limit); intervals (for each of average_queue_length, average_system_time, average_wait, loss_probability, server_utilization, throughput_per_time_unit: mean, lower, upper, half_width, relative_half_width, or status: "unavailable"); spread_across_replications (per metric min, max); t_critical; precision (target, rows of metric, relative_half_width, replications_needed, and negligible_wait_excluded); benchmark (method, applies - true only for a steady-state study -, values, within_interval); windows (count, window_length, queue_length_mean, warm_up_window, first_post_warm_up, later_average); comparison (null, or method, changes, pairs, scenario_b, benchmark_b, intervals_b, rows of metric, a_mean, b_mean, diff_mean, diff_lower, diff_upper, excludes_zero, better, and browser_call); metrics (M1.., each metric, value, sometimes lower / upper, a basis, and fields such as per_replication, metric_name, replications_needed, target, within_interval or excludes_zero); flags (F1.. with severity high / medium / low, category, message and refs); browser_verdict; expected (the exact values a SimPy reproduction must match, at full precision, including the two _seed_rep0 seeds as digit strings) and expected_count; and clipped.

A trimmed but real facts object for the page's walk-in vaccination clinic example (2 nurses, a waiting room of 20, a customer every 4 minutes, a 6-minute shot, 480 minutes, 20 replications, scenario B with 3 nurses). Entries shown as "..." are cut here for length; the page computes and sends the full object:

{
  "settings": {
    "time_unit": "minutes",
    "target_relative_half_width": 0.05,
    "simpy_version": "4.1.2",
    "schema_version": "1.1",
    "seed_rule": "derive_seed(base_seed, replication, 'arrivals'|'service') = BLAKE2b digest_size 8 of 'simpy-skill-v1|{base_seed}|{replication}|{stream}', big-endian"
  },
  "config": {
    "model": {
      "analysis_mode": "terminating", "horizon": 480, "warm_up": 0,
      "servers": 2, "queue_capacity": 20, "mean_interarrival": 4, "mean_service": 6,
      "base_seed": "20260723", "max_entities": 10000, "max_events": 200000
    },
    "replications": 20,
    "confidence": 0.95
  },
  "offered": {"arrival_rate": 0.25, "service_rate": 0.166667, "offered_load": 1.5, "rho_per_server": 0.75, "system_capacity": 22},
  "totals": {"arrivals": 2492, "rejected": 0, "admitted": 2492, "completed": 2404, "observed_completed": 2404,
             "unfinished": 88, "replications_with_unfinished": 20, "entity_limit_hits": 0, "no_completions": 0,
             "max_events_used": 832, "event_limit": 200000},
  "intervals": {
    "average_wait": {"mean": 6.40839, "lower": 4.78338, "upper": 8.0334, "half_width": 1.62501, "relative_half_width": 0.2536},
    "average_queue_length": {"mean": 1.70105, "lower": 1.25627, "upper": 2.14583, "half_width": 0.444783, "relative_half_width": 0.2615},
    "server_utilization": {"mean": 0.758245, "lower": 0.710961, "upper": 0.805528, "half_width": 0.0472836, "relative_half_width": 0.06236},
    "...": "average_system_time, loss_probability, throughput_per_time_unit"
  },
  "spread_across_replications": {"average_wait": {"min": 0.792059, "max": 15.0574}, "...": "5 more"},
  "t_critical": 2.0930241,
  "precision": {
    "target": 0.05,
    "rows": [
      {"metric": "average_queue_length", "relative_half_width": 0.2615, "replications_needed": 547},
      {"metric": "average_wait", "relative_half_width": 0.2536, "replications_needed": 515},
      "... 3 more"
    ],
    "negligible_wait_excluded": false
  },
  "benchmark": {
    "method": "M/M/c/K steady state (exponential arrivals and service, c servers, capacity c + queue_capacity)",
    "applies": false,
    "values": {"average_wait": 7.58296, "loss_probability": 0.00051044, "...": "7 more"},
    "within_interval": {"average_wait": true, "loss_probability": false, "...": "4 more"}
  },
  "windows": {"count": 20, "window_length": 24, "queue_length_mean": [0.347, 0.7081, 0.8365, "... 17 more"],
              "warm_up_window": 0, "first_post_warm_up": 0.347, "later_average": 1.684},
  "comparison": {
    "method": "paired differences B - A across replications with common random numbers",
    "changes": {"servers": {"a": 2, "b": 3}},
    "pairs": 20,
    "scenario_b": {"model": {"servers": 3, "...": "as config.model"}, "replications": 20, "confidence": 0.95},
    "benchmark_b": {"...": "as benchmark.values, for B"},
    "intervals_b": {"...": "as intervals, for B"},
    "rows": [
      {"metric": "average_wait", "a_mean": 6.40839, "b_mean": 0.924186, "diff_mean": -5.4842,
       "diff_lower": -6.90669, "diff_upper": -4.06171, "excludes_zero": true, "better": "b"},
      {"metric": "server_utilization", "a_mean": 0.758245, "b_mean": 0.513716, "diff_mean": -0.244529,
       "diff_lower": -0.259383, "diff_upper": -0.229675, "excludes_zero": true, "better": "not_judged"},
      "... 4 more"
    ],
    "browser_call": "b_better"
  },
  "metrics": [
    {"id": "M1", "metric": "offered load per server (rho)", "value": 0.75,
     "basis": "mean_service / (servers x mean_interarrival) = 6 / (2 x 4); at or above 1 the queue can only be held by its cap of 20"},
    {"id": "M4", "metric": "average wait in queue", "value": 6.40839, "lower": 4.78338, "upper": 8.0334,
     "half_width": 1.62501, "relative_half_width": 0.2536,
     "basis": "95% Student-t interval across 20 replications, completed customers who arrived at or after warm-up"},
    {"id": "M9", "metric": "worst relative half-width", "value": 0.2615, "metric_name": "average_queue_length",
     "replications_needed": 547, "target": 0.05, "basis": "..."},
    "... M2, M3, M5-M8, M10-M14"
  ],
  "flags": [
    {"id": "F1", "severity": "low", "category": "censoring",
     "message": "88 admitted customers (3.53% of those admitted and tracked) were still waiting or in service at the horizon and are left out of the wait averages, which biases them low only slightly.",
     "refs": "M8"},
    {"id": "F2", "severity": "medium", "category": "precision",
     "message": "The average queue length interval is +/-26.1% of its mean, wider than the 5% target; about 547 replications would reach it.",
     "refs": "M9"}
  ],
  "browser_verdict": "caveated",
  "expected": {
    "arrival_seed_rep0": "4022215348497587693",
    "service_seed_rep0": "14857552695425099618",
    "total_arrivals": 2492,
    "total_rejected": 0,
    "total_completed": 2404,
    "t_critical": 2.0930240544083016,
    "mean_average_wait": 6.4083885867886154,
    "half_width_average_wait": 1.6250106068402077,
    "diff_mean_average_wait": -5.4842023119996774,
    "...": "12 more: mean_ and half_width_ of every interval metric, diff_mean_server_utilization, diff_mean_loss_probability"
  },
  "expected_count": 21,
  "clipped": []
}

The replication counters and intervals in this object are what the skill's own replication_runner.py prints for the same configuration and seed (the page's Report .json (runner layout) button saves that report, and config.json (replication_runner.py) saves the config to pass it with --config). So over the API you can build facts yourself from the runner's output, adding the derived parts (offered, precision, benchmark, windows, comparison, metrics, flags, browser_verdict, expected), or you can simply use the page, which computes all of it for free before any spend. Send the object as a string: json.dumps(facts), JSON.stringify(facts) or your language's equivalent. The script lane copies expected into its reproduction check, so keep those values at full precision.

Building the body

The simplest way to get a body that matches the page byte for byte is to run the page's own modules in Node. deskkit.js needs only simkit.js (the SimPy port) next to it, and both export themselves with module.exports. The form values are strings, exactly as the page's inputs hold them; blank fields take the runner's defaults, and any b_ field (b_servers, b_queue_capacity, b_mean_interarrival, b_mean_service) adds scenario B.

// make-body.js - build the exact body the page sends, with the page's own code.
// Save https://sim-desk.skillsafe.ai/deskkit.js and https://sim-desk.skillsafe.ai/simkit.js next to
// this file, edit the form values below, then:   node make-body.js > body.json
const K = require("./deskkit.js");
const set = {
  lane: "interpret", title: "Walk-in vaccination clinic", time_unit: "minutes",
  context: "Walk-in vaccination clinic open 8 hours. ... I think two nurses keep the average wait under 5 minutes.",
  question: "Is a third nurse worth it?", decision: "",
  servers: "2", queue_capacity: "20", mean_interarrival: "4", mean_service: "6", horizon: "480",
  analysis_mode: "terminating", warm_up: "0", base_seed: "20260723", max_entities: "10000", max_events: "200000",
  replications: "20", confidence: "0.95",
  b_servers: "3",                           // scenario B: leave every b_ field out for no comparison
};
const pl = K.plan(set);                     // validates exactly as replication_runner.py does
if (!pl.ok) throw new Error(pl.errors.join("; "));
const a = K.runAll(pl.a), b = pl.b ? K.runAll(pl.b) : null;
const X = K.analyze(set, pl, a, b);
const body = K.mustBeObject(K.buildInput(X, set));
console.error("browser verdict:", X.hint, "| flags:", X.flags.map(f => f.id + " " + f.category).join(", "));
console.error("idempotency key: sim-desk:" + body.task + ":" + K.hashInput(body) + ":a1");
console.log(JSON.stringify(body));
# Or build the body in any language from a facts object you already hold (for example one you
# assembled from replication_runner.py's report). facts must go in as a STRING.
#   python3 build_body.py facts.json > body.json
import json, sys

facts = json.load(open(sys.argv[1]))
body = {
    "task": "interpret",
    "title": "Walk-in vaccination clinic",
    "context": "Walk-in vaccination clinic open 8 hours. ... I think two nurses keep the average wait under 5 minutes.",
    "question": "Is a third nurse worth it?",
    "facts": json.dumps(facts, separators=(",", ":")),
}
print(json.dumps(body))

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"sim-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your study.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="sim-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://sim-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"sim-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://sim-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"sim-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra; markup_bps is the app's markup in basis points. hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to interpret or script, facts non-empty, and facts a JSON string that parses to an object (the page never spends without an object body, DeskKit.mustBeObject, and a facts string built by DeskKit.buildInput).

# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("interpret","script") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("facts",)) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, sim-desk:<lane>:<hash>:a<attempt> (for example sim-desk:interpret:d88f90a3f179d873:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: a changed configuration, seed, replication count, scenario B or notes (so changed facts) or a changed reading are a new hash, the same study in the other lane is a new key, and replaying an old key with a different body is a 409. The page uses DeskKit.hashInput(body) for the hash (it covers task, title, context, facts, decision and question; make-body.js prints the key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # interpret or script
KEY="sim-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"interpret\",\"verdict\":\"caveated\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])'

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"interpret\",\"verdict\":\"caveated\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is a JSON object serialised as a string. Parse it, then check the lane.

# The reply is a JSON string inside data.output.output. Pull it out and parse it:
printf '%s' "$OUT" | python3 -c 'import sys,json;r=json.loads(json.load(sys.stdin)["data"]["output"]["output"]);print(r["verdict"],r["headline"])'

Invariants worth asserting

The page reconciles every reply against the facts before it shows it (its recon.js, loadable in Node like deskkit.js). A direct API caller gets the model's reply as it is, so do your own checks; these are the ones the page makes:

The output contract

Every key of the lane's contract is present; empty sections are []. Strings are plain sentences under 500 characters (the script under 12,000).

{
  "lane": "interpret" | "script",
  "verdict": "sound" | "caveated" | "unreliable",
  "headline": "one sentence",
  "tldr": ["2-5 bullets; one starts \"Answer:\" when a question was asked"],
  // interpret:
  "metrics": [{"id": "M1", "reading": "..."}],          // one per facts.metrics item, same order
  "study": ["1-3 strings: terminating or steady state, horizon and warm-up, replications, what the intervals cover"],
  "comparison": {"call": "b_better|a_better|trade_off|no_clear_difference|not_run", "reading": "..."},
  "claims": [{"claim": "...", "support": "supported|partly|not_supported", "why": "..."}],
  "cautions": ["1-4 strings on what the study cannot show"],
  // script:
  "fixes": [{"fix": "...", "why": "...", "refs": "F2"}],  // 1-5; refs "" for a plain reproduction step
  "script": "import hashlib\nimport math\nimport random\nimport simpy\n...",
  "assumptions": ["1-4 strings"],
  "checks": ["1-4 strings"],
  // both:
  "next_steps": ["1-5 concrete actions"],
  "prescan_responses": [{"ref": "F1", "verdict": "confirmed|dismissed", "note": "..."}]
}

The verdict rule: high flags set unreliable, medium flags set caveated, low flags never change it. comparison.call is copied from the browser, and a scenario is called better on a metric only where that row's paired interval excludes zero (lower wait, time in system, queue length and loss are better, higher throughput is better, utilization is neither). support is supported when the relevant interval lies wholly on the claimed side, not_supported when it lies wholly on the other side, and partly when it straddles.

The script lane's script follows a fixed order: the configuration constants from facts.config (BASE_SEED as an int); EXPECTED; derive_seed(base_seed, replication, stream) as the BLAKE2b digest of f"simpy-skill-v1|{base_seed}|{replication}|{stream}"; run_replication(cfg, replication) with separate "arrivals" and "service" random.Random streams, a simpy.Resource of servers, rejection when servers + queue_capacity places are taken, and the queue and utilization integrated over [warm_up, horizon); replications 0 to replications - 1 with half_width = T_CRITICAL * stdev / sqrt(n) and T_CRITICAL = EXPECTED["t_critical"]; scenario B with the same seed when a comparison was run; every key checked with == (integers) or math.isclose(got, want, rel_tol=1e-9, abs_tol=1e-12); then the fixes, printing the new intervals; and finally the replication configuration as JSON for replication_runner.py --config. A value you did not give is a named constant set to None with an AUTHOR_INPUT_NEEDED comment, and that fix is skipped while it is None.

Worked example: interpret

The page's walk-in vaccination clinic example. With 2 nurses the browser finds an average wait of 6.41 minutes (4.78 to 8.03), a queue-length interval of +/-26.1% of its mean against a 5% target (F2, medium) and 88 customers censored at the horizon (F1, low), so its read is caveated; scenario B with 3 nurses is b_better. The body below is trimmed: the facts string is cut down for readability, and a real request sends the full string from make-body.js or the page:

{
 "task": "interpret",
 "title": "Walk-in vaccination clinic",
 "context": "Walk-in vaccination clinic open 8 hours. People arrive about every 4 minutes and a shot takes about 6 minutes; the waiting room holds 20. We want to know whether a third nurse is worth it. I think two nurses keep the average wait under 5 minutes.",
 "question": "Is a third nurse worth it?",
 "facts": "{\"settings\":{\"time_unit\":\"minutes\",\"target_relative_half_width\":0.05,\"simpy_version\":\"4.1.2\",\"...\":\"more\"},\"config\":{\"model\":{\"analysis_mode\":\"terminating\",\"horizon\":480,\"warm_up\":0,\"servers\":2,\"queue_capacity\":20,\"mean_interarrival\":4,\"mean_service\":6,\"base_seed\":\"20260723\",\"max_entities\":10000,\"max_events\":200000},\"replications\":20,\"confidence\":0.95},\"offered\":{\"arrival_rate\":0.25,\"service_rate\":0.166667,\"offered_load\":1.5,\"rho_per_server\":0.75,\"system_capacity\":22},\"totals\":{\"arrivals\":2492,\"rejected\":0,\"completed\":2404,\"unfinished\":88,\"...\":\"more\"},\"intervals\":{\"average_wait\":{\"mean\":6.40839,\"lower\":4.78338,\"upper\":8.0334,\"half_width\":1.62501,\"relative_half_width\":0.2536},\"...\":\"5 more\"},\"...\":\"spread_across_replications, t_critical, precision, benchmark, windows, metrics M1-M14 as above\",\"comparison\":{\"changes\":{\"servers\":{\"a\":2,\"b\":3}},\"pairs\":20,\"rows\":[\"... 6 rows\"],\"browser_call\":\"b_better\"},\"flags\":[{\"id\":\"F1\",\"severity\":\"low\",\"category\":\"censoring\",\"message\":\"88 admitted customers (3.53% of those admitted and tracked) were still waiting or in service at the horizon and are left out of the wait averages, which biases them low only slightly.\",\"refs\":\"M8\"},{\"id\":\"F2\",\"severity\":\"medium\",\"category\":\"precision\",\"message\":\"The average queue length interval is +/-26.1% of its mean, wider than the 5% target; about 547 replications would reach it.\",\"refs\":\"M9\"}],\"browser_verdict\":\"caveated\",\"expected\":{\"arrival_seed_rep0\":\"4022215348497587693\",\"total_arrivals\":2492,\"...\":\"19 more\"},\"expected_count\":21,\"clipped\":[]}"
}

An excerpt of the real reply (arrays cut where they show "..."):

{
 "lane": "interpret",
 "verdict": "caveated",
 "headline": "In this terminating 480-minute, 2-nurse clinic model average wait is 6.41 minutes and queue length 1.70 with a wide (about 26%) precision margin, while a paired 3-nurse scenario shows meaningfully shorter waits, queues and system time at lower utilization (b_better).",
 "tldr": [
  "Answer: In this model, moving from 2 to 3 nurses (scenario B) lowers average wait, queue length and time in system, each with a paired interval that excludes zero (browser_call b_better), at the cost of lower server utilization (75.8% to 51.4%); this is a within-model comparison, not a real-world guarantee.",
  "With 2 nurses, average wait is estimated at 6.41 minutes (95% CI 4.78 to 8.03), above the user's 5-minute target, though the CI's lower edge does dip under it.",
  "..."
 ],
 "metrics": [
  {
   "id": "M4",
   "reading": "Those same customers waited an average of 6.41 minutes before being seen (95% CI 4.78 to 8.03 minutes)."
  },
  "... M1-M3, M5-M14"
 ],
 "study": [
  "This is a terminating-horizon study: all 20 replications start empty and run for exactly 480 minutes with no warm-up (warm_up = 0), so the results describe this single 8-hour opening, not a long-run average.",
  "..."
 ],
 "comparison": {
  "call": "b_better",
  "reading": "Scenario B (3 servers) shows lower average queue length, average system time and average wait, and higher throughput, each with a paired 95% interval that excludes zero; server utilization is also lower with a non-zero interval but is not itself better or worse, and loss probability shows no clear difference (both estimated at 0)."
 },
 "claims": [
  {
   "claim": "Two nurses keep the average wait under 5 minutes.",
   "support": "partly",
   "why": "The average_wait point estimate for the 2-server model is 6.41 minutes, above 5, but its 95% CI (4.78 to 8.03) spans both sides of 5 minutes, so the facts lean against the claim without ruling it out at the current precision."
  }
 ],
 "cautions": [
  "This model assumes exponential (memoryless) interarrival and service times and a single pool of interchangeable nurses; it does not show how real patients or a specific clinic layout would behave under other distributions.",
  "..."
 ],
 "next_steps": [
  "If tighter precision on average queue length is needed, run about 547 replications (precision.rows), since the current interval is about +/-26.1% of the mean.",
  "..."
 ],
 "prescan_responses": [
  {
   "ref": "F1",
   "verdict": "confirmed",
   "note": "totals.unfinished is 88 across the 20 replications, matching the flag's count of admitted-but-not-completed customers, and these are indeed excluded from the wait and time-in-system averages, so the concern applies."
  },
  {
   "ref": "F2",
   "verdict": "confirmed",
   "note": "intervals.average_queue_length.relative_half_width is 0.2615 (about 26.1%), matching the flag, well above precision.target of 0.05, so the concern applies."
  }
 ]
}

This reply answers F1 and F2 once each, stays at the browser's caveated, copies the call b_better, reads M1 to M14 in order, and judges the claim in context ("two nurses keep the average wait under 5 minutes") as partly, because the average_wait interval straddles 5.

Worked example: script

The same study with the interpret reading handed over as decision (the page fills it with Recon.decisionText of the earlier reply; it may be empty). The facts string is the same study's, sent in full. The body, with facts and decision trimmed:

{
 "task": "script",
 "title": "Walk-in vaccination clinic",
 "context": "Walk-in vaccination clinic open 8 hours. People arrive about every 4 minutes and a shot takes about 6 minutes; the waiting room holds 20. We want to know whether a third nurse is worth it. I think two nurses keep the average wait under 5 minutes.",
 "decision": "Verdict: caveated.\nIn this terminating 480-minute, 2-nurse clinic model average wait is 6.41 minutes ...\n- M1: ...\n- M14: ...\nComparison: b_better. ...\nNext steps:\n- ...",
 "facts": "{... the same study's facts string ...}"
}

The reply carries the common envelope plus fixes (here typically one answering F2 by running the 547 replications from precision.rows, with refs "F2"), script, assumptions and checks. The script must set BASE_SEED = 20260723, put all 21 keys of facts.expected into EXPECTED (arrival_seed_rep0 4022215348497587693, service_seed_rep0 14857552695425099618, total_arrivals 2492, total_rejected 0, total_completed 2404, t_critical 2.0930240544083016, the mean and half-width of all six interval metrics, and the three diff_mean_ values for scenario B), check each one, then apply the fixes. Run it with pip install simpy==4.1.2 and python sim_check.py; a clean reproduction prints no mismatch.

Truncation and partial results

If your balance sits between min_credits and hold_credits, the run still executes with a smaller output cap and the job carries "truncated": true. The JSON may then stop mid-object: close it (the page's Recon.closeJson does this) and show the sections that arrived, saying how many of the lane's sections were recovered, rather than treating a clipped reply as complete. A clipped script is not runnable; re-run instead.