# oneshot

**Receipts for one-prompt claims.**

"One prompt, one agent" is now a headline. Readers can't check it: the prompt is withheld, failed attempts are not
counted, and only the successes are shown. `oneshot` records every attempt an agent makes on a task — the prompt,
the model, each try, what was checked, which result was kept — in a hash-chained receipt that anyone can verify.
From the receipt comes an honest label, a disclosure score, and a check of the headline claim.

One file, zero dependencies, Node 18+ and the browser. MIT licence.

```
npm test                              # 13 tests
node bin/oneshot.js demo              # a toy solver on 40 exactly checkable problems
node bin/oneshot.js demo --edit-after 2 --attempts 5
```

## What a receipt proves

- **Every attempt is on it.** Each event carries the SHA-256 of the one before. Drop a failed attempt, change a token
  count or swap an output, and `verify()` names the event that no longer matches.
- **Which prompt was used, even before you show it.** The prompt is stored as a salted commitment. Publish the
  receipt without the prompt today (`redact()`); reveal it later (`reveal()`), and anyone can check it is the same
  prompt, word for word. A changed prompt mid-run is a new commitment, so it shows.
- **What the label is allowed to say.** `label()` reads the receipt and says one of:
  - *One prompt, one attempt*: a true one-shot.
  - *One prompt, best of N*: it took N tries; the receipt keeps all of them.
  - *Prompt changed K times*: someone edited the prompt between tries.
  - *No checked result*.
  Plus whether the kept result was checked, and how (`exact`, `lean`, `tests`, `human`).
- **What was left out.** `summarize()` over many receipts gives tasks tried vs results shown, one-shots vs best-of-N
  vs edited, and the attempts and tokens spent. `checkClaim()` tests four headline claims: `one-shot`,
  `one-prompt`, `all-checked`, `nothing-hidden`.
- **What a reader still needs.** `disclosure()` scores seven items: the prompt, the model, a reasoning summary, time,
  compute or cost, every attempt, and how the result was checked.

## Use

```js
const O = require("oneshot");              // or <script src="oneshot.js"> → window.OneShot

const run = O.createRun({ task: "double a number", model: "your-model", params: { temperature: 0.7 }, prompt });
for (let k = 1; k <= 5; k++) {
  const out = await callModel(prompt), ok = runTests(out);
  run.attempt({ output: out, ok, check: "tests", tokens, seconds });   // failures too
  if (ok) break;
}
const receipt = run.close();               // or run.close({ redact: true }) to publish without the prompt

O.verify(receipt)        // { ok: true, problems: [], id }
O.label(receipt).text    // "One prompt, best of 3"
O.badge(receipt)         // an SVG badge for your README
```

From the command line, pipe one JSON object per attempt:

```
printf '{"output":"draft","ok":false,"check":"tests","tokens":900,"seconds":11}\n{"output":"final","ok":true,"check":"tests","tokens":700,"seconds":8}\n' \
  | node bin/oneshot.js record --task "Fix the parser" --prompt "Fix the failing test." --model local-7b > receipt.json
node bin/oneshot.js label receipt.json     # One prompt, best of 2 · checked (tests) · disclosure 6/7
```

## API

| | |
|---|---|
| `createRun({ task, prompt, model, params, by, salt })` | `attempt({ output, ok, check, tokens, seconds, cost, summary, prompt })`, `keep(tryNo)`, `close({ redact })` |
| `verify(receipt)` | chain, totals, outputs and revealed prompts: `{ ok, at, problems, id }` |
| `label(receipt)` | `{ code, grade, text, checked, attempts, prompts, kept }` |
| `disclosure(receipt)` | `{ score, of: 7, items }` |
| `summarize(receipts)`, `checkClaim(receipts, claim)` | many runs at once |
| `redact(receipt)`, `reveal(receipt, prompts, salt)` | publish now, show the prompt later |
| `replay(receipt, solver)` | run it again with the revealed prompts and compare every output |
| `badge(receipt)` | SVG |
| `toy.problems / solve / check / model / experiment` | a small, exactly checkable problem set and a pretend model |

## What the toy shows (60 problems, seed 7)

| setting | results | one attempt | best of N | after a prompt edit | no result | attempts |
|---|---|---|---|---|---|---|
| 1 attempt, temperature 0 | 28 | 28 | 0 | 0 | 32 | 60 |
| 3 attempts, temperature 0 | 28 | 28 | 0 | 0 | 32 | 124 |
| 1 attempt, temperature 0.7 | 31 | 31 | 0 | 0 | 29 | 60 |
| 3 attempts, temperature 0.7 | 47 | 31 | 16 | 0 | 13 | 105 |
| 5 attempts, edit after 2 fails | 58 | 31 | 13 | 14 | 2 | 119 |

The headline "58 of 60 solved" and the honest "31 solved in one attempt with one prompt" describe the same run.
Retries cost twice the attempts and only help when the model samples (temperature above 0).

## Honest limits

- A receipt proves what was recorded, not that everything was recorded. It is only as complete as the loop that
  writes it; run the loop where others can see it, or have a third party hold the run.
- Timestamps come from the machine that records them.
- A checked result is only as good as its check. "exact" in the toy recomputes the answer; a proof checker confirms
  the proof of the statement as written, not that the statement is the right one.

## Licence

MIT
