← Pairwise Desk / API
Tokens

Drive Pairwise Desk from your own code

Everything the web page does is available over HTTP: send the Microsoft PICT model, say what the feature does, attach the facts and the covering array you generated in your own process, and get the same test designer back — a judgement on whether the model is the right model for the feature, or a named, runnable test case for every generated row. The natural use is a pull request that changes a model and wants it reviewed, or a nightly job that re-names the suite whenever the spec moves.

One thing to be clear about before the first call: the model never computes anything. Parsing the PICT text, generating the t-wise covering array, proving that every achievable tuple is covered and that every row satisfies every constraint, counting the excluded tuples and the lower bound — all of that is done by the caller and sent as facts. The language model does judgement over those facts: whether the parameters are the parameters the feature actually has, whether a constraint is too strict, which error path deserves a negative value, what each row is really a witness for. See computing the facts yourself; the engine the web page uses ships as one plain script, /pict.js, that loads in node.

And one thing to be clear about the suite: a t-wise covering array is a greedy construction. It is proven complete against its own coverage target on every run — that proof is facts.self_check, which the engine recomputes each time — but it is not proven minimal, and Microsoft PICT itself may produce a suite of a different size and a different order from the same model. The desk generates tests; it does not run them.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{ "ok": true,  "data":  { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "status": 402, "details": { ... } } }

The token is minted for this app (the guest endpoint takes {"slug":"pairwise-desk"} in its body), so no slug header is needed afterwards — no app-slug header exists on this API at all. Send your token as Authorization: Bearer … on every call.

The run body is the input object: post {"task": "review", …} directly. Wrapping it as {"input": {…}} returns 200 and quietly hides every field from the model, so never do that.

StatusCodeMeaning
400validation_errorThe body is not a JSON object, a declared required field (task, spec, model, facts) is missing, or a field is not in the table below. The body is scalar-only, which is why facts and rows travel as JSON strings rather than objects or arrays.
401unauthorizedNo token, or a stale one. Mint a guest token or sign in again.
402insufficient_creditsThe balance is under min_credits. Price with /estimate first.
403forbiddenA guest token tried to run: running is metered and needs a personal token. Sign in to run.
404not_foundUnknown job id.
429rate_limitedBack off and retry.
5xxserver_errorTransient. Retry with the same Idempotency-Key so a retry never double-bills.

1. Get a token

A guest token is free and enough for /me and /estimate. Running a lane is metered, so it needs a personal token — sign in on the token page and copy it from there. A guest token that calls /run gets 403.

curl -s -X POST https://api.skillsafe.ai/v1/app-api/guest -H "Content-Type: application/json" -d '{"slug":"pairwise-desk"}'

2. A tiny client

One helper, one envelope. The samples below reuse it.

curl -s -X GET https://api.skillsafe.ai/v1/app-api/me \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json"

3. Check the session and the balance

GET /me returns subject_type (user or guest), subject_id and credits — and nothing else, so a signed-in caller is exactly subject_type === "user". A real user's first run should not 402: compare credits with the estimate's hold_credits before running.

curl -s -X GET https://api.skillsafe.ai/v1/app-api/me \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json"

4. Price the run — free

POST /estimate with the exact run body returns model (gpt-5.6-terra), model_alias (gpt-terra), markup_bps (1000), hold_credits, min_credits and sponsor_enabled. hold_credits is a reservation placed against the balance while the job runs, not the price: the unused part is refunded and the real cost comes back as charged_credits on the finished job. No job is created and nothing is charged by estimating, and a guest token may call it.

curl -s -X POST https://api.skillsafe.ai/v1/app-api/estimate \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"task": "review", "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.", "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";", "facts": "<JSON string from Pict.analyze - see below>"}'

The task field, and the two lanes

task is the first field to decide, and the reply always names the lane it answered in lane. The two contracts are never blended: one request gets one lane.

taskWhat it answersNeeds rows
review Is this the right model for the feature the spec describes? The designer reads the parameters, the values, the negatives and the constraints against the spec and proposes operations — add a parameter, add or remove a value, mark a value negative, add or remove a constraint, rename a parameter — then judges every constraint in the model one by one. Verdict: complete, gaps, rework or insufficient_spec. no
cases Turn every generated row into a test case a tester can run: a title, what the row is a witness for, two to five steps and the expected result, each derived from that row’s actual values and the spec. A row holding a ~ value is a negative test. Verdict: named, partial or insufficient_spec. yes

Any other value of task is treated as unrecognised: the closest lane is chosen from the input and named in lane.

The fields both lanes take

FieldTypeMeaning
taskstring, requiredreview or cases, as above.
specstring, requiredWhat the feature or product does, in your own words. It may be short; everything the designer says about what should happen is grounded in it, and where it is silent the reply says so rather than guessing. In the cases lane it is what the expected results are derived from.
modelstring, requiredThe PICT model text as written: one Name: v1, v2 line per parameter, optional ~negative values, a | alias aliases, (n) weights, { A, B } @ n sub-models, then the IF … THEN …; constraints. A very long model may be clipped in the middle with a bracketed line saying so; the facts are always computed from the whole text, never from the clipped copy.
factsstring, requiredA JSON string (not an object), produced by Pict.analyze(model, {order: 2}) with rows, self_check and summary.value_counts removed, and excluded capped at forty entries. Everything the designer is allowed to quote as a number comes from here. See computing the facts yourself.
rowsstringcases lane only, and required there: a JSON string holding the array of test rows to name, in order — [{"id": 1, "UserType": "Guest", …}, …], exactly facts.rows, capped at the first thirty. A cell shown as ~Expired is a negative value and makes that row a negative test. Every row gets exactly one case with the same id; nothing is added, dropped, merged or renumbered.
retry_notestringOptional, and normally absent. The page sends it only on the automatic reformat retry, when the first reply did not parse as one JSON object.

Every field is a scalar string. Anything not in this table is rejected as unknown, and a missing task, spec, model or facts is a 400 validation_error.

Computing the facts yourself

Load the one engine script in node with a stub window and call the same function the page calls. Everything the designer is allowed to quote comes out of this one call:

global.window = {};
const vm = require("vm"), fs = require("fs");
vm.runInThisContext(fs.readFileSync("pict.js", "utf8"));   // the one file: https://pairwise-desk.skillsafe.ai/pict.js
const Pict = window.Pict;

const model = [
  "UserType: Guest, Registered, Premium",
  "PaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired",
  "ShippingMethod: Standard, Express, Pickup",
  "Region: US, EU, APAC",
  "IF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";"
].join("\n");

const facts = Pict.analyze(model, { order: 2 });   // order 2 is pairwise; 3..6 raise the strength

// facts.order              the t actually generated (clamped to the parameter count)
// facts.summary            parameters, values, constraints, submodels, order,
//                          tests, product, reduction_pct,
//                          tuples_required, tuples_covered, tuples_excluded, coverage_pct,
//                          lower_bound, negative_values, negative_rows,
//                          largest_parameter, unused_values[], value_counts{}
// facts.params[]           {name, values[], negative[], aliases[], weighted[], count}
// facts.constraints[]      each constraint as the parser read it back, one string each
// facts.submodels[]        {params[], order}
// facts.rows[]             {id, <Parameter>: <value>, ..} - the suite; a ~ prefix marks a negative value
// facts.excluded[]         {cells: [[param, value], ..], scope} - tuples the constraints make unachievable
// facts.self_check         {rows_valid, invalid_rows[], duplicates[], required_tuples, covered_tuples,
//                           complete, negative_rows, rows_with_two_negatives} - the proof, recomputed every
//                           run and kept local: it is your evidence, not something the designer re-checks
// facts.warnings[]         engine warnings the reply must address in notes_on_input
// facts.errors[]           what could not be parsed at all; a non-empty errors means there is no suite

const lean = JSON.parse(JSON.stringify(facts));
delete lean.rows;                    // the rows travel in their own field, and only in the cases lane
delete lean.self_check;              // the proof stays with you; nothing in the contract quotes it
delete lean.summary.value_counts;    // per-value tallies: bulky, and nothing in the contract quotes them
if (lean.excluded.length > 40) {     // the page sends the first forty and says how many it held back
  lean.excluded_omitted = lean.excluded.length - 40;
  lean.excluded = lean.excluded.slice(0, 40);
}

JSON.stringify(lean)                      // <- send this string as the facts field
JSON.stringify(facts.rows.slice(0, 30))   // <- send this string as the rows field, cases lane only

Pict.toCsv(facts)                         // the suite as CSV
Pict.toMarkdown(facts)                    // the suite as a Markdown table
Pict.applyOperations(model, operations)   // {text, applied[], rejected[], errors[]} - apply a review reply

For the model above the engine returns twelve tests against a product of 108 combinations, sixty-two required tuples all covered, one tuple excluded (a guest paying by bank transfer, which the constraint forbids) and a lower bound of twelve — so this suite happens to be as small as any pairwise suite over this model can be, and facts.warnings says so. Those are the numbers the reply quotes; it derives none of its own.

Send JSON.stringify(lean) as facts — a string, not an object — and, in the cases lane, JSON.stringify(facts.rows.slice(0, 30)) as rows. The thirty-row cap is the web page’s: a longer suite is named in batches, each batch its own run, and the ids stay the engine’s ids so the cases line up with the rows they came from. The generator is deterministic — the same model text yields the same suite — so the same facts and the same rows can be regenerated at any time and an Idempotency-Key derived from them is stable.

The review request body, in full

Every field the page sends in the review lane, with facts abbreviated:

{
  "task": "review",
  "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.",
  "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";",
  "facts": "<JSON string from Pict.analyze - see below>"
}

Worked example: the review lane

Four parameters, one negative value and one constraint, against a checkout the spec describes in three sentences. Price it first; the same body goes to /run.

curl -s -X POST https://api.skillsafe.ai/v1/app-api/estimate \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"task": "review", "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.", "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";", "facts": "<JSON string from Pict.analyze - see below>"}'

The cases request body, in full

The same four fields plus rows, the array the designer must name one case per row:

{
  "task": "cases",
  "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.",
  "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";",
  "facts": "<JSON string from Pict.analyze - see below>",
  "rows": "<JSON string from facts.rows.slice(0, 30) - see below>"
}

Worked example: the cases lane

The same model and spec, now asking for a runnable case per generated row. rows is the engine’s own array, ids and all.

curl -s -X POST https://api.skillsafe.ai/v1/app-api/estimate \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"task": "cases", "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.", "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";", "facts": "<JSON string from Pict.analyze - see below>", "rows": "<JSON string from facts.rows.slice(0, 30) - see below>"}'

5. Run it, then poll

POST /run returns {job_id}; GET /jobs/{job_id} until status is succeeded or failed. Send an Idempotency-Key header derived from the lane, the model text and the facts so a retry never double-bills — the engine is deterministic, so that key is stable across regenerations. The reply’s output.output is the JSON text described in the contract below, charged_credits is the actual cost once the unused part of the hold is released, and truncated is true when the balance cut the output short.

curl -s -X POST https://api.skillsafe.ai/v1/app-api/run \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: pairwise-desk-review-0001" \
  -d '{"task": "review", "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.", "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";", "facts": "<JSON string from Pict.analyze - see below>"}'
curl -s -X GET https://api.skillsafe.ai/v1/app-api/jobs/JOB_ID \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json"

6. Or stream it

POST /run-stream is server-sent events: job with the job id, then tick heartbeats, then done carrying the finished job. Browsers receive ticks rather than text deltas, so build progress on elapsed time and parse the output from done. The cases lane is the long one — it writes a title, steps and an expected result for every row — so it is the one worth streaming.

curl -N -s -X POST https://api.skillsafe.ai/v1/app-api/run-stream \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Accept: text/event-stream" \
  -d '{"task": "cases", "spec": "Checkout for a web store. A shopper checks out as a guest or signed in, pays by card, PayPal or bank transfer, picks a shipping method and buys from one of three regions. An expired card must be refused at the payment step with a message that names the card, and a guest cannot pay by bank transfer because there is no account to invoice.", "model": "UserType: Guest, Registered, Premium\nPaymentMethod: CreditCard, PayPal, BankTransfer, ~Expired\nShippingMethod: Standard, Express, Pickup\nRegion: US, EU, APAC\nIF [UserType] = \"Guest\" THEN [PaymentMethod] <> \"BankTransfer\";", "facts": "<JSON string from Pict.analyze - see below>", "rows": "<JSON string from facts.rows.slice(0, 30) - see below>"}'
# events: job (the job id), tick (heartbeat), done (the finished job). Browsers receive ticks, not deltas.

The output contract

One JSON object, no prose around it and no code fences, no Markdown inside the strings. Common keys on every reply, whichever lane answered:

KeyTypeMeaning
lanereview | casesThe lane actually answered. If task was missing or unrecognised the closer lane is chosen from the input and named here.
titlestringShort, naming the feature the model is for.
headlinestringOne sentence: the verdict and the one thing that decides it.
verdictstringOne of the lane’s enum, below.
summarystringTwo or three sentences, readable instead of the rest.
notes_on_input[]string[]Anything unclear, missing or contradictory in the spec or the model, including every entry of facts.warnings.
risks[]string[]What could go wrong if the suite is run as it stands.
next_steps[]string[]Ordered and concrete.

An empty list is [], never omitted. The review lane answers with this shape:

{ "lane": "review",
  "title": "short title naming the feature",
  "headline": "one sentence: the verdict and the biggest gap or strength",
  "verdict": "complete" | "gaps" | "rework" | "insufficient_spec",
  "reading": "3-6 sentences. Must quote facts.summary.tests (the number of tests) and facts.summary.parameters (the number of parameters) exactly. Says what the suite covers, what the constraints exclude and whether that matches the spec, and what is missing.",
  "operations": [
    { "op": "add_parameter", "name": "NewParam", "values": ["a", "b", "~bad"], "why": "one sentence tied to the spec" },
    { "op": "add_value", "parameter": "Existing", "value": "c", "negative": false, "why": "..." },
    { "op": "remove_value", "parameter": "Existing", "value": "a", "why": "..." },
    { "op": "mark_negative", "parameter": "Existing", "value": "a", "why": "..." },
    { "op": "add_constraint", "pict": "IF [A] = \"x\" THEN [B] <> \"y\";", "why": "..." },
    { "op": "remove_constraint", "pict": "<verbatim text from facts.constraints>", "why": "..." },
    { "op": "rename_parameter", "parameter": "Old", "name": "New", "why": "..." }
  ],
  "constraint_review": [ { "constraint": "<verbatim from facts.constraints>", "verdict": "keeps" | "too_strict" | "too_loose" | "redundant", "note": "one sentence" } ],
  "coverage_notes": [ "what pairwise coverage over this model does and does not prove for this feature; 1-4 items" ],
  "notes_on_input": [ "anything unclear, missing or contradictory in the spec or model" ],
  "risks": [ "what could go wrong if the suite is run as is" ],
  "next_steps": [ "ordered, concrete" ],
  "summary": "2-3 sentences for a status update" }

At most twelve operations. Each names a parameter that exists — except add_parameter, which must name one that does not; a value already listed is never added and the only value of a parameter is never removed; an add_constraint must be valid PICT, [Param] in square brackets, string values in double quotes, IF … THEN …; or a bare predicate ending in ;, operators = <> < <= > >= IN { } LIKE combined with AND OR NOT. A verdict of complete carries zero operations; gaps and rework carry at least one. Every constraint in facts.constraints is judged exactly once in constraint_review, quoted verbatim.

The cases lane answers with this shape:

{ "lane": "cases",
  "title": "short title naming the feature and the suite",
  "headline": "one sentence on what the suite exercises and where the spec leaves the expected results open",
  "verdict": "named" | "partial" | "insufficient_spec",
  "cases": [ { "id": 1,
               "title": "imperative, under 80 characters, naming the distinguishing values",
               "kind": "positive" | "negative",
               "exercises": "the pairs or behaviour this row is the witness for, in one sentence",
               "steps": "2-5 numbered steps in one string, separated by ' / '",
               "expected": "the observable result, specific to this row" } ],
  "gaps": [ "behaviours the spec describes that no row exercises, or expected results the spec leaves undefined" ],
  "notes_on_input": [ ], "risks": [ ], "next_steps": [ ],
  "summary": "2-3 sentences" }

cases holds one entry per row in rows, in the same order and with the same id: nothing added, dropped, merged or renumbered. A row whose cell shows a ~ value is a negative test — its kind is negative and its expected is the error or rejection the spec describes for that value. Where the spec does not say what should happen, the expected result is written in the row’s own parameter terms and the silence is listed in gaps; that is what partial means. A case never names a value the row does not hold.

How the browser reconciles the reply

Nothing in the reply is taken on trust, and a pipeline of your own should do the same three checks the web page does. Numbers: every figure the designer writes is meant to be a verbatim copy out of facts, so pull the numbers back out of reading, headline and each why, compare them with the engine, and surface the mismatches rather than the prose — a test count, a percentage or a parameter count that is not in facts is a disagreement, not a finding. Operations: the browser does not paste a rewritten model, it applies each operation with Pict.applyOperations(model, operations), re-parses and re-generates, and shows the before and the after. An operation the engine rejects — a parameter that does not exist, a value already listed, a constraint that will not parse — comes back in rejected with the reason and is shown as a disagreement rather than silently dropped. Ids: in the cases lane the set of cases[].id must equal the set of rows[].id exactly, and each case must name only values its own row holds; a missing id, an extra id or a value from a different row is caught before anything is displayed. The regenerated suite from an accepted set of operations is then the suite the cases lane is handed, which is how a review turns into named tests without anyone retyping the model.

Derived from the agent skill @sickn33/pypict-skill — the PICT Test Designer workflow, whose upstream is omkamal/pypict-claude-skill — for the test-design judgement both lanes follow. The model language and its semantics are Microsoft PICT, microsoft/pict; the generator here is an independent implementation of t-wise covering-array generation over that language, so its suite may differ from PICT’s in size and in order while covering the same tuples. The suite is greedy, proven complete against its own coverage target on every run and not proven minimal. The desk designs tests; it does not run them. Not affiliated with the skill’s author or with Microsoft.