EVIDENCE FIRST LOCAL PROTOTYPE · 01

AI WORKFLOW / AGENT ENGINEERING

A correct number.
An incomplete answer.

An agent computes 140. The real engineering question is whether it knows what that number leaves out.

A reproducible local tool workflow that binds results to sources, preserves analytical definitions and keeps release decisions explicit.

Synthetic data. Actual Python execution. No live model benchmark.

RUN / demo-002SYNTHETIC

Known revenue subtotal

140units
Known contributions4 of 5 rows

One missing price. A complete revenue total is unknown.

Tool executionCompleted
Analytical releaseBlocked
Both states can be true.
01Strict tool requests
02Source-bound results
03Explicit failure states
04Reproducible evidence

01 / SYSTEM DESIGN

Separate reasoning from execution.
Keep the evidence between them.

A small, inspectable boundary between a host that requests work and a local executor that can prove what it did.

01 / READ

Persistent contract

Purpose, definitions, authority and current state survive the conversation.

AGENTS.md + contracts/
02 / REQUEST

Host + agent

Inspect inputs, then request a supported tool with the current source digest.

JSON in · fixed tools
03 / EXECUTE

Local executor

Validate arguments, calculate, check eligibility and retain artifacts.

Python · standard library
04 / VERIFY

Evidence + handoff

Recheck source and result integrity. Carry limitations into the next step.

Result + receipt
SEPARATE AUTHORITY

Native application review and human release approval remain outside this executor. A successful run cannot grant either.

Why a narrow runner?

The prototype accepts three named tools through a strict JSON contract. It does not turn model-generated text into a shell command. That makes behavior and failure paths easier to inspect.

What is the security boundary?

The host and operating system still control permissions and isolation. Path checks and an allowlist are application controls. This project does not claim to be an OS sandbox.

02 / FAILURE EXPLORER

Change the condition.
Inspect what the agent may claim.

This browser illustration explains the contract. Selecting a scenario does not execute Python or call a model.

RECORDED VALUES / EXPLANATORY VIEWRELEASE BLOCKED

Report the subtotal. Preserve the limits.

The arithmetic is valid for the known contributions. Missing price and duplicated customer keys restrict the answer.

Instruction access
Readable files; host loading unverified
Tool execution
Local Python completed
Evidence status
Recorded receipt; rerun to assess current files
Release decision
Blocked for complete revenue and customer join
PERMITTED HANDOFF

Known subtotal: 140 fictional units across 4 of 5 rows. Complete total unknown. Customer-segment join blocked by duplicate C2.

Next action: resolve the missing price and customer-key conflict through the data owner; retain the current qualified result.

03 / ANALYTICAL CONTRACT

Small data. Consequential distinctions.

The fixture is small enough to recompute by hand. These semantics are preserved before an agent interprets the result.

DENOMINATOR

91 / 110 = 82.7273%

Pool successes and eligible cases. Averaging the group rates, 10% and 90%, would produce 50%.

MEANING OF ZERO

Unknown ≠ observed zero

Filling a blank price with zero can leave the displayed subtotal unchanged while making a false completeness claim.

Missingdistinct states0

JOIN CARDINALITY

Duplicate C2 blocks the join

A repeated dimension key can multiply transaction rows. Deduplication requires an explicit rule, not a silent repair.

C2
Segment BSegment C
Inspect all five synthetic transactions
Fictional currency units. Negative quantity represents a return.
TransactionQuantityUnit priceDiscountContribution
T1210010%180
T2−110010%−90
T31500%50
T41Missing0%Unknown
T5100%0

Known subtotal: 180 − 90 + 50 + 0 = 140. Four of five rows are known. No revenue-coverage percentage or causal conclusion follows.

04 / EXECUTION EVIDENCE

Inspect the receipt.
Reproduce the result.

This section presents a retained local run. It does not check your current filesystem or certify a model host.

LOCAL EXECUTION

Verification report included

Read the retained logs for the exact test scope and limitations.

Runtime
Python standard library
Inference calls
None
Native review
Not performed
Host compatibility
Not established
Download verification report
REPRODUCE / FROM PROJECT ROOT
python -B examples/run_demo.py --workspace . --output-dir runs/my-review

Use a new output directory for each demonstration. Python 3.10+ is the intended baseline; see the report for the version actually tested.

RETAINED RESPONSEDownload JSON ↓
Read the actual tool response
Loading bundled receipt...

05 / ENGINEERING JUDGMENT

Choose controls for the failure.
Keep the limits visible.

Portable definitions

Persistent rules and hash-bound inputs travel with the project. Actual loading, tool calling and permissions still depend on the host.

Useful integrity checks

Checksums can detect changed bytes. They do not authenticate the author, establish source truth or protect records from an actor who controls them all.

Proportionate complexity

A standard-library runner keeps this example inspectable. It does not provide centralized lineage, production concurrency or infrastructure isolation.

Native capability preserved

Excel/BI and managed plugins can belong in the broader workflow. Their availability and native acceptance must be demonstrated in the actual application.

Practices behind the design

Specific comparisons, not a claim of industry certification.

AQuA · assurancedbt · data assertionsOpenLineage · run identityAnthropic · agent evaluation

NEXT VERIFIED MILESTONE

Connect one host.
Measure actual behavior.

The next step is a named Claude or open-weight execution host, a fresh-session loading check and a bounded evaluation plan. This prototype makes that integration reviewable; it does not claim it has happened.

Read the host integration procedure