← DESIGN.md Desk / API
Tokens

Drive DESIGN.md Desk from your own code

The lint is exact; the judgement is the model's. The model reads your DESIGN.md and the official linter's findings; it never re-lints and is told never to compute a contrast ratio. Re-lint a repaired file yourself (the page does) before you commit it.

Everything the web page does is available over HTTP. Lint the file with the page's own dmdmd.js and dmdkit.js (the official @google/design.md 0.4.0 linter, built unmodified for the browser) and deskscan.js, send the file and the facts, and get back a verdict (blocked, needs_work, ready) and either a review or a repaired file. The natural loop: lint, review, repair, re-lint, review again. DESIGN.md Desk is derived from the agent skill @nousresearch/design-md (Hermes Agent).

Two lanes: the task field

taskwhat comes backextra input
reviewAn answer to every finding and check, prose that disagrees with the tokens, gaps, what a coding agent would get wrong, 3-7 priority fixes.none
repairThe complete repaired DESIGN.md, every change against the finding it answers, open questions for the author.review: the review to apply (optional)

An unknown or missing task is answered as review, and the headline says so.

Input fields

Every field is a string. task, design_md and facts are required.

fieldwhat it holds
taskreview or repair.
design_mdThe DESIGN.md text. The page sends up to 40,000 characters; past that the front matter goes whole and the middle of the prose is cut on line boundaries with a marker line.
factsA JSON string (see below).
titleOptional name for the run, up to 160 characters.
contextOptional notes: who reads the file, what worries you. Up to 3,000 characters.
questionOptional. When present, the first next_steps entry starts with Answer:.
reviewRepair only: the review to apply, as text. The page builds it from a review reply.
retry_noteOnly on a reformat retry, telling the model what was wrong with its last reply.

The facts string

{
  "linter": "@google/design.md 0.4.0 (official, run in the browser)",
  "spec_version": "alpha",
  "name": "Quayvane",
  "summary": {"errors": 1, "warnings": 5, "infos": 2},
  "findings": [{"id": "F1", "severity": "error", "rule": "broken-ref", "path": "components.status-chip",
                "message": "Reference {colors.accent} does not resolve to any defined token."}, ...],
  "checks": [{"id": "P1", "kind": "prose-hex-not-token", "text": "The prose in ## Colors writes #0B5FFF, ..."}],
  "tokens": {"colors": {"primary": "#0f2a3d", ...}, "typography": {...}, "rounded": {...}, "spacing": {}, "components": {...}},
  "contrast": [{"component": "button-secondary", "background": "#f4f6f8", "text": "#8a94a6", "ratio": 2.82, "passes_aa": false}, ...],
  "sections": {"present": ["Overview", "Typography", "Colors", ...], "canonical_order": [...], "missing": ["Elevation & Depth", "Shapes"], "unknown": []},
  "hint": "blocked",
  "clipped": []
}

findings are the linter's, in its order; checks are the page's own four prose checks (prose-hex-not-token, prose-ref-unresolved, role-undocumented, component-undocumented). hint is blocked when an error stands, needs_work when a warning or check stands, otherwise ready. From another language you can build findings from npx @google/design.md lint DESIGN.md (number them F1, F2 ...) and leave checks empty; the page's own build is exact.

Building the body

The shortest exact path is the page's own modules in Node. Download dmdmd.js, dmdkit.js and deskscan.js next to this script:

// make-body.mjs - node make-body.mjs DESIGN.md review > body.json
import { readFileSync } from "node:fs";
import vm from "node:vm";

const ctx = { console };
ctx.globalThis = ctx; ctx.window = ctx;
vm.createContext(ctx);
for (const f of ["dmdmd.js", "dmdkit.js", "deskscan.js"]) vm.runInContext(readFileSync(f, "utf8"), ctx);
const D = ctx.DeskScan;

const [file = "DESIGN.md", task = "review"] = process.argv.slice(2);
const scan = D.scan(readFileSync(file, "utf8"));
const body = D.buildInput(scan, { lane: task, title: "", context: "", question: "", review: "" });
console.error("lint:", JSON.stringify(scan.summary), "hint:", scan.hint,
  "key: design-md-desk:" + task + ":" + D.hashInput(body) + ":a1");
console.log(JSON.stringify(body));

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"design-md-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your text.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="design-md-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://design-md-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"design-md-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://design-md-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"design-md-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra and markup_bps 1000 (a 10% markup). hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to review or repair, design_md and facts non-empty, and facts a JSON string that parses to an object (this is what the page's own guard, DeskScan.mustBeObject, refuses to spend without).

# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.mjs above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("review","repair") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("design_md", "facts")) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, design-md-desk:<lane>:<hash>:a<attempt> (for example design-md-desk:review:mt6z7s12xnl3e:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: an edited DESIGN.md (so changed facts), changed notes or a changed review are a new hash, the same file in the other lane is a new key, and replaying an old key with a different body is a 409. The page uses DeskScan.hashInput(body) for the hash (it covers task, design_md, facts, title, context, review and question; make-body.mjs prints the key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # review or repair
KEY="design-md-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"review\",\"verdict\":\"needs_work\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"review\",\"verdict\":\"needs_work\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is one JSON object as a string. Strip optional code fences, take the text from the first { to the last } and parse it. If it does not parse, send the same body once more with retry_note explaining the problem and the next attempt number in the Idempotency-Key - that is one extra, named run. The page's own parser is Recon.parseResult and Recon.normalize in recon.js.

Invariants worth asserting

The output contract

{
  "lane": "review" | "repair",
  "verdict": "blocked" | "needs_work" | "ready",
  "headline": "...",
  "finding_responses": [{"ref": "F1", "action": "fix" | "accept" | "dispute", "note": "..."}],
  // review:
  "prose_vs_tokens": [{"where": "## Colors", "issue": "...", "fix": "..."}],
  "gaps": [{"area": "colors|typography|layout|shapes|elevation|components|states|accessibility|prose", "issue": "...", "suggestion": "..."}],
  "agent_readiness": ["..."],
  "priority_fixes": ["...", "...", "..."],
  // repair:
  "changes": [{"ref": "F1", "change": "..."}],
  "design_md": "---\nversion: alpha\n...",
  "open_questions": ["..."],
  "next_steps": ["..."]
}

Worked example: review

The Quayvane example on the page: a broken {colors.accent} reference, a 2.82:1 secondary button, a misspelled spaceing map. Body (abridged):

{
  "task": "review",
  "design_md": "---\nversion: alpha\nname: Quayvane\ndescription: A calm, high-density operations console for freight dispatchers who watch it all shift.\ncolors:\n  primary: \"#0F2A3D\"\n  secondary: \"#... (the whole file)",
  "facts": "{\"linter\":\"@google/design.md 0.4.0 ...\",\"summary\":{\"errors\": 1, \"warnings\": 5, \"infos\": 2}, \"findings\":[...], ...}",
  "title": "Quayvane dispatch console",
  "context": "We hand this file to our coding agents before any UI work. Our Tailwind export seems to ignore the spacing scale.",
  "question": "What will an agent get wrong first?",
  "review": ""
}

Reply (abridged; the page replays the full reply for free):

{
  "lane": "review",
  "verdict": "blocked",
  "headline": "Quayvane is blocked by one broken reference (status-chip's {colors.accent}) and has a misspelled 'spaceing' key that is silently dropping the entire xs/sm/md/lg spacing scale from exports; fix the reference first, the typo second.",
  "finding_responses": [
    {
      "ref": "F1",
      "action": "fix",
      "note": "{colors.accent} in components.status-chip does not resolve to any token, so status-chip has no valid backgroundColor. Point it at an existing token, and per the components-list description (on time, at risk, late) split status-chip into per-state sibling components (e.g. status-chip-on-time, status-chip-at-risk, status-chip-late) each with a real color, since a state is a separate sibling component, not a single component with one ambiguous color."
    },
    {
      "ref": "F2",
      "action": "fix",
      "note": "components.button-secondary's textColor #8a94a6 on backgroundColor #f4f6f8 is 2.82:1, below the 4.5:1 AA minimum. Darken the text or change the background until the pair clears 4.5:1."
    },
    "..."
  ],
  "prose_vs_tokens": [
    {
      "where": "## Colors",
      "issue": "Prose gives Tertiary as #0B5FFF; colors.tertiary resolves to #0a5cff, a different hex.",
      "fix": "Change the prose hex to #0a5cff, or write {colors.tertiary} instead of a literal hex."
    },
    "..."
  ],
  "gaps": [
    {
      "area": "states",
      "issue": "The Components section says status-chip marks a load as on time, at risk or late, but only one status-chip component is defined and its backgroundColor ({colors.accent}) doesn't resolve. There are no status-chip-on-time / status-chip-at-risk / status-chip-late siblings, so an agent has no way to render the three states with distinct colors.",
      "suggestion": "Define status-chip-on-time, status-chip-at-risk and status-chip-late as separate sibling components (per the spec, states are siblings, not a nested map), each with a resolvable backgroundColor."
    },
    "..."
  ],
  "agent_readiness": [
    "It will render every status-chip with no visible background because {colors.accent} doesn't resolve to anything (F1), leaving on-time/at-risk/late loads visually identical and unstyled.",
    "It will ship button-secondary with text at 2.82:1 contrast against its background, well under the 4.5:1 AA minimum, making cancel/back buttons hard to read (F2)."
  ],
  "priority_fixes": [
    "Fix components.status-chip.backgroundColor: point {colors.accent} at an existing or new documented token instead of an unresolved reference (F1).",
    "Rename the 'spaceing' key to 'spacing' so the xs/sm/md/lg scale is recognized and stops being dropped by exports (F7, F8).",
    "Fix components.button-secondary's contrast: change textColor or backgroundColor so the pair clears 4.5:1 AA (F2).",
    "..."
  ],
  "next_steps": [
    "Answer: The status-chip component's backgroundColor reference ({colors.accent}) doesn't resolve, so an agent will render every status chip unstyled before it even gets to spacing or contrast problems. Separately, your Tailwind export is losing the spacing scale because the front matter key is spelled 'spaceing' instead of 'spacing', not because of an export bug.",
    "Rename 'spaceing' to 'spacing' in the front matter."
  ]
}

Worked example: repair

The same file, with the review above handed over as review. Body (abridged):

{
  "task": "repair",
  "design_md": "---\nversion: alpha\nname: Quayvane\ndescription: A calm, high-density operations console for freight dispatchers who watch it all shift.\ncolors:\n  primary: \"#0F2A3D\"\n  secondary: \"#... (the whole file)",
  "facts": "{\"linter\":\"@google/design.md 0.4.0 ...\",\"summary\":{\"errors\": 1, \"warnings\": 5, \"infos\": 2}, \"findings\":[...], ...}",
  "title": "Quayvane dispatch console",
  "context": "We hand this file to our coding agents before any UI work. Our Tailwind export seems to ignore the spacing scale.",
  "question": "",
  "review": "Review verdict: blocked. Quayvane is blocked by one broken reference (status-chip's {colors.accent}) and has a misspelled 'spaceing' key that is silently dropp..."
}

Reply (abridged):

{
  "lane": "repair",
  "verdict": "needs_work",
  "headline": "Quayvane's blocking reference and every warning are fixed by reusing tokens already in the file, but the three new status-chip variants haven't been contrast-checked yet, so this still needs one more re-lint before it's ready.",
  "finding_responses": [
    {
      "ref": "F1",
      "action": "fix",
      "note": "components.status-chip.backgroundColor referenced {colors.accent}, which does not exist. A single chip also can't show three different states with one colour, so I replaced it with three sibling components (components.status-chip-on-time, components.status-chip-at-risk, components.status-chip-late), each with its own resolvable backgroundColor from tokens already in the file."
    },
    {
      "ref": "F2",
      "action": "fix",
      "note": "components.button-secondary's textColor #8a94a6 on backgroundColor #f4f6f8 was 2.82:1, below AA. Changed it to backgroundColor {colors.surface} / textColor {colors.on-surface}, the exact pair components.table-row already uses at a verified 14.8:1."
    },
    "..."
  ],
  "changes": [
    {
      "ref": "F1",
      "change": "Removed components.status-chip (broken backgroundColor {colors.accent}) and added components.status-chip-on-time ({colors.secondary} / {colors.surface}), components.status-chip-at-risk ({colors.brand-sand} / {colors.primary}) and components.status-chip-late ({colors.error} / {colors.surface}), each keeping the original typography and rounded values."
    },
    {
      "ref": "F2",
      "change": "components.button-secondary: backgroundColor changed from {colors.neutral} to {colors.surface}, textColor changed from {colors.muted} to {colors.on-surface}."
    },
    {
      "ref": "F3",
      "change": "colors.brand-sand is now referenced by components.status-chip-at-risk and documented in ## Colors."
    },
    "..."
  ],
  "design_md": "---\nversion: alpha\nname: Quayvane\ndescription: A calm, high-density operations console for freight dispatchers who watch it all shift.\ncolors:\n  primary: \"#0F2A3D\"\n  secondary: \"#4A6275\"\n  tertiary: \"#0A5CFF\"\n  neutral:... (the whole repaired file)",
  "open_questions": [
    "status-chip-on-time uses colors.secondary and status-chip-at-risk uses colors.brand-sand as stand-in state colours, since the palette has no dedicated green/amber token for 'on schedule' or 'caution'; confirm these read correctly or supply dedicated tokens.",
    "The three new status-chip variants have not been through the browser's contrast check; confirm colors.secondary/{colors.surface}, colors.brand-sand/{colors.primary} and colors.error/{colors.surface} each clear 4.5:1 once re-linted."
  ],
  "next_steps": [
    "Re-lint the repaired file and confirm the three new status-chip variants pass AA contrast.",
    "Confirm or replace the stand-in colours chosen for status-chip-on-time and status-chip-at-risk."
  ]
}

Truncation and partial results

If your balance sits between min_credits and hold_credits, the run still executes with a smaller output cap and the job says "truncated": true. The JSON may then be cut mid-object: close the open strings and brackets, parse what arrived, and treat missing keys as missing - never as empty. The page does this with Recon.closeJson and says how many sections were recovered. A repair cut short has no usable design_md; run it again.

DESIGN.md Desk is derived from the agent skill @nousresearch/design-md. The linter is @google/design.md 0.4.0 (Apache-2.0, Copyright 2026 Google LLC; notice).