# Evidence-first agent workflow

A reproducible local tool prototype and portfolio case study for AI workflow and agent engineering. It demonstrates how an agent-facing execution boundary can preserve analytical definitions, reject stale context and report success without falsely approving an incomplete result.

**Start with [the offline site](dist/index.html), [the case study](docs/CASE_STUDY.md), or [the actual verification record](docs/VERIFICATION.md).** All business data are synthetic. No model account or inference service is needed.

## What works

- Three executable tools: `inspect_inputs`, `analyze`, `verify_result`.
- Strict JSON requests, a current-source digest precondition and structured failures.
- A source-bound result, receipt and checksum record in a new run directory.
- Exact decimal revenue arithmetic, pooled conversion, explicit unknowns and customer-join checks.
- Behavioral regression tests, a runnable example and a portable static presentation.

This is an application-level prototype, not an OS sandbox or production agent service. Claude Code and open-weight integration are documented but not executed here. Native Excel/BI review and human release approval remain separate.

## Reproduce in three steps

Use a local Python interpreter. Python 3.12.14 was used for the retained execution; 3.10+ is the intended syntax baseline, not a tested version matrix. No third-party Python dependencies are required. Commands below run from this project directory. If `python` is not on PATH, substitute the path to an existing interpreter. Do not install anything just to read the site.

```text
python -B -m unittest discover -s tests -v
python -B examples/run_demo.py --workspace . --output-dir runs/my-review
python -B src/tool_runner.py --workspace . --output-dir runs/my-review
```

The first command creates and removes its own disposable fixtures under `evidence/engine/test-workspaces/`. The second performs inspection, analysis and verification, then prints all responses. It creates `runs/my-review/demo-001/`. The third starts the one-request CLI; supply the following JSON on stdin and close stdin to verify the previous run:

```json
{"request_id":"review-001","tool":"verify_result","arguments":{"run_id":"demo-001"}}
```

Use a fresh output directory for another demonstration. An existing run is not overwritten. On PowerShell, a one-line JSON string can be piped directly:

```powershell
'{"request_id":"review-001","tool":"verify_result","arguments":{"run_id":"demo-001"}}' | python -B src/tool_runner.py --workspace . --output-dir runs/my-review
```

Read [the protocol](contracts/tool-protocol.md) for exact fields, exit codes, path rules, cancellation and retries. The tool response includes verification limits; preserve them when summarizing it.

## Expected analytical result

Overall conversion is 91/110, approximately 82.7273%. The known revenue subtotal is 140 fictional currency units across four of five rows. One price is unknown. A zero price is observed, and a negative quantity remains a return. Duplicate customer C2 blocks the customer-segment join. Complete revenue and customer-join release remain ineligible even though the computation succeeds.

## Project map

| Path | Purpose |
|---|---|
| `dist/index.html` | Offline case study and illustrative failure explorer |
| `src/tool_runner.py` | Real local JSON executor |
| `contracts/` | Executable metric definitions and tool protocol |
| `data/` | Small synthetic CSV inputs |
| `examples/run_demo.py` | Three-call reproduction procedure |
| `tests/` | Behavioral tests using actual subprocess invocations |
| `evidence/engine/` | Retained test logs and real run artifacts |
| `docs/VERIFICATION.md` | Current test scope and unresolved limits |
| `docs/DECISIONS.md` | Architecture, trust model and trade-offs |
| `docs/HOSTS.md` | Host integration procedure and capability boundaries |
| `docs/PRACTICES.md` | Comparison with specific primary-source practices |
| `docs/INTERVIEW.md` | Five-minute demonstration and learning exercises |
| `scripts/build_portfolio.py` | Rebuild static evidence assets and shareable archives |

## View and rebuild the site

Open `dist/index.html` directly. Its assets are local, and it does not fetch remote fonts, scripts or data. Links to external reference documentation require internet access. The Copy command button may need a secure browser context; its fallback explains how to copy manually.

For a normal local website preview:

```text
python -B -m http.server 8765 --bind 127.0.0.1 --directory dist
```

Open `http://127.0.0.1:8765/`. Stop the process with Ctrl+C when finished. This serves only static files from `dist`; it does not expose the executor as a network service.

After an intentional source or evidence change, rerun the relevant checks and refresh the verification report before rebuilding:

```text
python -B scripts/build_portfolio.py
python -B scripts/check_portfolio.py
```

Additional retained checks use `python -B scripts/check_reproduction.py` and `python -B scripts/check_download.py`. A Node.js runtime is optional for `node scripts/check_site.cjs`, which checks interaction logic with a minimal DOM test double and does not render the site. Use `python -B tests/run_suite.py` when intentionally refreshing the hash-bound behavioral test report, then review the actual outcomes before rebuilding.

The build copies selected documentation and retained evidence into site assets. It creates a source archive for site visitors and a full review archive in `release/`. These are local files; no publication occurs. Do not rebuild a passing badge around changed, unverified code.

## Scope, authorship and rollback

This is a user-directed, AI-assisted learning project. It does not claim real client impact, professional deployment or independently authored implementation. [Attribution and sharing notes](RIGHTS.md) distinguish public-safe content from permission to publish or choose a code license.

The project is independent of the original analytics workspace and its sealed candidate. Its commands do not modify those projects. Git and Docker were excluded. No credentials, plugin permissions or security settings were changed.

To roll back, stop this project's preview process and archive or remove this independent directory and its generated archives. For a single demo run, remove only the output directory you created after verifying it belongs to that run. Preserve retained evidence if you want to reproduce the portfolio's reported results.
