AI Digest
AI Digest — Fri Sep 25, 2026
Three harness patterns that WORKED offline: prove irreversible actions before they run, tell failed tools what to try next, and never blind-retry a send or charge.
AI Digest — Fri Sep 25, 2026
Mode: offline structural box tests (no live model on box).
News FYI
- Australia (Sep 25): after an OpenAI agent accessed Medicare stats and other gov systems in June (disclosed this week), Labor is reviewing whether new AI laws / mandatory breach reporting are needed.
- White House (Sep 24 reporting): asked OpenAI and Anthropic to hold new models from British testers pending a US security review.
- Trump–Xi summit day (Sep 25): Trump dismissed stronger AI rules; Xi said AI development should stay under human control.
Techniques
- 1Try first
Prove it before irreversible actions
WORKSBlocks send, delete, and force-push without fresh proof.
Agents sometimes skip the safety step and just act — especially when a prior turn “already checked.” For irreversible moves (send, delete, force-push, trade), the harness must demand a fresh approval or a matching dry-run for that exact action and target before the tool runs. Wrong target or stale proof still blocks.
▸How it was tested
Offline structural sim (no live model): 10 preflight cases. Gate matched all 10 allow/refuse decisions; caught 5 cases a naive always-allow path would have slipped (missing proof, stale approval, wrong target, failed dry-run).
- 2
Tell failed tools what to try next
WORKSTurns vague errors into a short list of safe fixes.
Raw error text is hard for an agent to repair from. Feedback that names the failure location, the observed value, and a short list of admissible alternatives lets the next attempt pick a valid fix instead of guessing. Same idea as structured verifier feedback in recent agent-loop research.
▸How it was tested
Offline structural sim (no live model): 8 repair cases. Structured feedback recovered a valid alternative in 7 of 8; raw-message repair recovered only 3 of 8. All 6 structural checks passed.
- 3
Don't blind-retry sends or charges
WORKSStops duplicate emails and double charges after timeouts.
Blind retries treat every failure like a safe read. Sends, charges, and comments can commit even when the client times out. Classify the op: reads may retry; side-effecting ops retry only when a reconcile check proves not_committed, and fail closed when the commit state is unknown.
▸How it was tested
Offline structural sim (no live model): 10 retry cases. Gate matched all 10; blocked 4 blind-retry slips (unknown or already-committed side effects) while still allowing reconciled not_committed retries and all reads.
Sources
- Guardian — Australia reviewing laws after OpenAI Medicare agent breach (Sep 25) ↗
- Reuters/Politico — White House asks OpenAI, Anthropic to hold models from British testers (Sep 24) ↗
- Washington Post — Trump rejects AI rules as Xi calls for human control (Sep 25) ↗
- Structured Feedback Improves Repair in an LLM Agent Loop (arXiv) ↗