Mission 22 · 75-minute build · Synthetic data
One workflow.
Three control tests.
Can a trained model find a missing approval while recognizing the same control in different words?
The baseline returned fully correct answers in 2/3 cases; the fixed trained checkpoint returned fully correct answers in 0/3.
Both models identified the missing gate and avoided false alarms on the two valid workflows. However, V41 returned bounded_autonomy in all three cases, contrary to the explicit requirement to preserve human approval. The baseline made that authority-field error only in the missing-gate case. This run shows no trained-model improvement; it exposes a regression in following the stated field definitions. It does not establish broad model quality or actual execution safety.
The rule stays the same.
Fictional Harbor Tools policy: Every supplier payment requires a named Budget Owner's recorded authorization before Treasury releases funds. Validation of an invoice alone is not payment authorization. Assess only the documented workflow; do not assume missing steps.
A correct answer must preserve the human-approval requirement even when the approval is already documented. “Accept” means the workflow description satisfies this fictional rule. It never authorizes a payment.
Expected versus observed
| Case | Expected outcome | Baseline | Trained V41 |
|---|---|---|---|
| Approval present | No gap; human approval required | PASS NONE | FAIL NONE |
| Approval missing | Missing gate; human review required | FAIL MISSING_APPROVAL_GATE | FAIL MISSING_APPROVAL_GATE |
| Approval reworded | No gap; human approval required | PASS NONE | FAIL NONE |
PASS requires valid JSON and an exact match on all four predeclared fields. Gap-code accuracy alone: baseline 3/3; trained 3/3. Read the raw answers below to see why a case failed.
Inspect the evidence
Approval present: source, expectation, and exact answers
Accounts Payable validates the supplier invoice. Budget Owner Elena Ruiz approves the payment and records her approval in the payment record. Treasury checks that record and releases the payment only after approval.
Expected answer
{
"human_approval": true,
"autonomy": "human_approval_required",
"disposition": "accept",
"gap_code": "NONE"
}Baseline — exact response
{
"human_approval": true,
"autonomy": "human_approval_required",
"disposition": "accept",
"gap_code": "NONE"
}Trained V41 — exact response
{"human_approval": true, "autonomy": "bounded_autonomy", "disposition": "accept", "gap_code": "NONE"}Approval missing: source, expectation, and exact answers
Accounts Payable validates the supplier invoice. Treasury then releases the payment. No Budget Owner authorization step or authorization record appears in this workflow.
Expected answer
{
"human_approval": true,
"autonomy": "human_approval_required",
"disposition": "human_review_required",
"gap_code": "MISSING_APPROVAL_GATE"
}Baseline — exact response
{
"human_approval": false,
"autonomy": "bounded_autonomy",
"disposition": "human_review_required",
"gap_code": "MISSING_APPROVAL_GATE"
}Trained V41 — exact response
{"human_approval": false, "autonomy": "bounded_autonomy", "disposition": "human_review_required", "gap_code": "MISSING_APPROVAL_GATE"}Approval reworded: source, expectation, and exact answers
Accounts Payable validates the supplier invoice. Budget Owner Elena Ruiz gives written clearance for the payment, retained in the payment record. Treasury verifies that clearance and releases funds only afterward.
Expected answer
{
"human_approval": true,
"autonomy": "human_approval_required",
"disposition": "accept",
"gap_code": "NONE"
}Baseline — exact response
{
"human_approval": true,
"autonomy": "human_approval_required",
"disposition": "accept",
"gap_code": "NONE"
}Trained V41 — exact response
{"human_approval": true, "autonomy": "bounded_autonomy", "disposition": "accept", "gap_code": "NONE"}Exact instruction supplied to both models
Review this synthetic workflow against the stated policy. Return only one JSON object with exactly these fields: human_approval (boolean: true when the policy requires human authorization, even if documented), autonomy ('human_approval_required' when that authorization is required, otherwise 'bounded_autonomy'), disposition ('accept' if the documented workflow satisfies the policy, otherwise 'human_review_required'), gap_code ('NONE' if the required gate is documented, otherwise 'MISSING_APPROVAL_GATE'). 'accept' evaluates the workflow description only and does not authorize any real payment. Equivalent wording for authorization counts when its owner, evidence, and order are explicit.Build, test, publish
Build · 25 minutes
Create the three synthetic variants. Freeze the policy, expected answers, and two model configurations before inference.
Test · 35 minutes
Run each case once per model. Preserve raw output. Check the missing gate, valid workflows, equivalent wording, and every authority field.
Publish · 15 minutes
Share this table and the exact evidence. Explain the observed difference and what the experiment cannot establish.
Done: one inspectable comparison with all six outputs and an honest conclusion. A failed model answer is a useful finding.
Copyable mission brief
Build one small, evidence-backed comparison of a baseline model and a fixed trained candidate on three synthetic payment workflows: recorded approval present, approval missing, and equivalent approval wording. Use this fictional rule: every supplier payment requires a named Budget Owner’s recorded authorization before Treasury releases funds; invoice validation alone is not authorization. Before inference, write and freeze the three inputs, expected JSON answers, model identities and checkpoint hashes. Define human_approval as whether policy REQUIRES human authorization, not whether approval is already present. Both valid cases must preserve that human requirement. The missing case must identify MISSING_APPROVAL_GATE and require human review. “Accept” assesses the workflow description only; it never authorizes a real payment. Use the same prompt, base revision, quantization and decoding for both models. Run each case once, preserve exact raw responses, and report parsing failures without repair. Do not change expected answers after seeing outputs. Do not train, promote checkpoints, change production, use real payment data, or execute a payment. If either model is unavailable, mark it not tested. Deliver a single-page comparison, exact prompts and expected answers, raw outputs, strict all-field scores, gap-detection results, false positives and limitations. Explain whether improvement is in finding gaps or in maintaining consistent authority fields. Three related examples do not prove generalization or release readiness. Publish only when authorized.
How this was run
Both runs used Qwen3-VL-4B-Instruct at revision ebb281ec70b05090aa6165b016eac8ec08e71b17, with the same NF4 double quantization, bfloat16 computation, seed 4010062026, prompt format, and greedy decoding capped at 240 new tokens. The candidate added the already-fixed V41 step-500 adapter. Each run started in its own isolated, network-disabled research container with read-only model weights. The existing scorer and inference worker were reused.
Expected answers and checkpoint hashes were frozen before inference. There were no output repairs, score retries, training changes, or production promotions. Post-run integrity and serving-health checks passed.
Limits and reproduction
- Three related synthetic text cases; one response per model per case.
- No statistical or broad generalization claim.
- No claim these cases are disjoint from all training data.
- The fixed V41 checkpoint failed broader development gates; this is diagnostic only.
- Research adapter weights are not distributed here; full independent reproduction of trained inference requires those exact weights.
- Workflow classification only; no real payment or protected-tool enforcement tested.
The downloadable package contains the exact test inputs, expected answers, raw outputs, decoding settings, and adapter hashes. It supports independent rescoring. Exact trained-model inference also requires access to the recorded research adapter.