# Local tool protocol 1.0

Run from the project root with Python 3.10 or newer. There are no third-party packages.

```text
python -B examples/run_demo.py --workspace . --output-dir runs/demo --run-id demo-001
python -B -m unittest discover -s tests -v
```

The demo invokes the real runner three times and prints three response lines. Repeating its run ID in the same output directory fails with `OUTPUT_COLLISION`. Use a new run ID or a new output directory. Existing evidence is never overwritten.

For one tool call, start `python -B src/tool_runner.py --workspace . --output-dir runs/demo` and send one JSON document on standard input followed by EOF. The runner emits exactly one compact JSON response line on standard output. Standard error stays empty for handled failures. The operator selects paths as process options. A tool request cannot choose a file path, execute a command, load code, or call a network service.

## Requests

Every request has exactly `request_id`, `tool`, and `arguments`. Unknown fields, duplicate JSON keys, nonstandard constants, malformed JSON and incorrect argument types are rejected. The input limit is 65,536 characters. IDs have 1 to 64 ASCII letters, digits, underscores or hyphens and start with a letter or digit.

| Tool | Exact arguments | Effects |
| --- | --- | --- |
| `inspect_inputs` | `{}` | Read fixed inputs, validate schemas and definitions, return source digest and release blockers. No writes. |
| `analyze` | `{"expected_source_digest":"<64 lowercase hex characters>"}` | Require the digest from inspection, calculate results and commit a new run directory. |
| `verify_result` | `{"run_id":"demo-001"}` | Read the selected run, compare checksums, source bytes, definitions and engine bytes, and recompute the result. No writes. |

```json
{"request_id":"inspect-001","tool":"inspect_inputs","arguments":{}}
```

`examples/analyze-request.json` contains a valid example for the shipped fixture. Always obtain a new digest from inspection after editing source bytes. `verify_result` uses the operator's same output directory and the analysis request ID as `run_id`.

## Responses

Every response has the following keys. No response includes absolute machine paths.

| Key | Type and meaning |
| --- | --- |
| `protocol_version` | String, `1.0`. |
| `request_id` | Valid request ID or null when unavailable. |
| `ok` | Boolean, whether this tool invocation succeeded. This does not grant release eligibility or human approval. |
| `status` | `success`, `error`, or `cancelled`. |
| `tool` | Allowlisted tool name, or null when unavailable. |
| `result` | Tool-specific object on success, otherwise null. |
| `error` | Null on success; otherwise `{code, message, retryable}`. Messages omit sensitive local paths. |
| `artifacts` | Analysis-only paths `{result, receipt, checksums}`, relative to the selected workspace; otherwise `{}`. |
| `verification` | Explicit limits: native BI `not_run`, host cancellation `not_tested`, OS sandbox `not_claimed`, real-model quality `not_tested`, human review `not_performed`. |

Successful execution exits 0. Handled errors exit 1. An observed local cancellation marker exits 3. `OUTPUT_BUSY` is retryable after the other writer finishes or an operator investigates a stale lock. Other errors require corrected input or an explicit decision; the executor never retries automatically.

`inspect_inputs.result` contains `source_digest`, a `sources` map of relative input paths to SHA256, `definition_id`, `row_counts`, and `release_blockers`.

`analyze.result` contains:

- `definition_id`, `source_digest`, and `synthetic: true`.
- `conversion`: integer `successes` and `eligible`, `rate` as a decimal string rounded to 12 places with Decimal's half-even rule, or null for zero eligible, and `method: pooled_denominator`.
- `revenue`: `units: fictional_currency_units`, decimal-string `known_subtotal`, integer `known_rows` and `total_rows`, nullable decimal-string `complete_total`, and ordered per-transaction `contributions` with `transaction_id` and nullable decimal-string `net_revenue`.
- `customer_join`: `status` is `blocked` or `eligible`, sorted `duplicate_keys` and `unmatched_keys`, and `joined_output: not_produced`. This executor validates join cardinality; it does not produce segment metrics.
- `release`: boolean `eligible`, an ordered list of blockers, and `human_approval: not_granted` even when eligibility is true.

`verify_result.result` contains `run_id`, `integrity: verified`, `source_digest`, `release_eligible`, `human_approval: not_granted`, and the completed `checks` list. An integrity failure produces an error response and a nonzero exit.

The shipped calculation is 91/110, or 0.827272727273. The known revenue contributions are 180, -90, 50, unknown and 0. The subtotal is 140 fictional units, with four known contributions across five transactions. The missing price and duplicate customer C2 block release. Neither is repaired silently.

## Inputs and bounds

The fixed source set is `data/transactions.csv`, `data/groups.csv`, `data/customers.csv`, and `contracts/metric-definitions.json`. Exact UTF-8 bytes, including definitions and line endings, are hashed. The combined source digest is SHA256 of the canonical JSON source-hash map, with sorted keys, compact separators and a trailing newline. Editing whitespace changes the digest.

CSV headers are exact and ordered. Each source is limited to one MiB and 10,000 nonempty rectangular data rows. Numeric values use plain finite decimal notation, at most 12 integer digits and six fractional digits. Quantities are signed integers. Prices and counts are nonnegative; discounts are in [0, 1]; successes cannot exceed eligible. Transaction and group identifiers must be unique. Customer IDs may repeat so the duplicate defect can be reported without losing the valid independent computations. Only a blank price represents an unknown contribution.

The definition file must equal the supported `synthetic-v1` semantic contract. A semantic change requires an explicit implementation update. Decimal arithmetic uses precision 80 during computation. Money values remain strings rather than binary floats. No currency conversion, real financial inference or external population claim is made.

Saved result, receipt and checksum files have a separate four MiB read limit. This accommodates JSON expansion from a valid source near the input limit. A regression uses 10,000 transactions with 64-character IDs to verify that the executor can read back its own permitted output.

## Write boundary and integrity

Output directories must resolve within the selected workspace. Outputs cannot equal, contain or sit inside `data`, `src` or `contracts`. Existing symlinks and junctions are resolved before boundary checks. Input and saved-result paths must also resolve inside the workspace.

Analysis acquires `<output-dir>/.writer.lock` exclusively. It creates three files under a unique pending directory, flushes file contents, rechecks source digest and cancellation, then renames the directory to `<output-dir>/<request_id>`. Each run contains `result.json`, `receipt.json`, and `checksums.json`. Directory rename is the commit point. No result directory is exposed with only some of these files. Existing run IDs are refused. Normal cleanup only removes temporary files created by the current invocation and its own lock.

The receipt binds the source map, combined source digest, engine SHA256, result SHA256, protocol version, definition ID and run ID. The checksum map also binds receipt bytes. Verification compares these records with current files and a fresh computation. Checksums detect changes but are not signatures: a party able to replace code, sources and all records is outside this threat model. Keep an independently retained package checksum for stronger provenance.

`--cancel-file <workspace-relative-path>` is an optional operator-selected local marker. The executor checks it before work, before writes and just before commit. A pre-commit marker prevents a committed result. This is cooperative cancellation, not evidence that any host application cancelled or killed a process. A marker appearing after the final check can arrive after commit. Process interruption can leave a lock and pending files; automatic lock stealing and deletion are deliberately absent. After confirming no owner is active, an operator can inspect and archive those leftovers.

This is application-level path validation, not an OS security sandbox. It assumes a trusted local operator and no hostile concurrent filesystem mutation. It does not eliminate time-of-check/time-of-use races, provide directory-fsync crash durability, authenticate receipts, or make a model obey instructions.

## Error codes

`INVALID_OPTIONS`, `REQUEST_LIMIT`, `INVALID_JSON`, `INVALID_REQUEST`, `UNKNOWN_TOOL`, `INVALID_ARGUMENTS`, `MISSING_CONTEXT`, `MISSING_WORKSPACE`, `PATH_BOUNDARY`, `INVALID_OUTPUT`, `CANCELLED`, `MISSING_INPUT`, `INPUT_LIMIT`, `INVALID_DATA`, `INVALID_DEFINITIONS`, `STALE_CONTEXT`, `OUTPUT_BUSY`, `OUTPUT_COLLISION`, `MISSING_RESULT`, `RECEIPT_TAMPERED`, `RESULT_TAMPERED`, `INVALID_SAVED_RESULT`, `STALE_SOURCES`, `ENGINE_CHANGED`, `IO_OR_NUMERIC_ERROR`, and `INTERNAL_ERROR` are handled as structured failures. Error messages explain the correction or boundary without claiming successful execution.
