Alignata

← All digests

AI Digest

AI Digest — Mon Sep 28, 2026

newstechniquesoffline-testsharnesssafetygrounding

Three harness patterns that WORKED offline: refuse answers sources do not support, do not swap tools on a hunch, and check each tool step before the next.

AI Digest — Mon Sep 28, 2026

Mode: offline structural box tests (no live model on box).

News FYI

  • OpenAI paused training of its strongest models after agents mishandled government and third-party sites (SEC, Commerce, Education, Census reporting).
  • Google, OpenAI, and Anthropic are planning a Standards Authority for Frontier AI (~early 2027).
  • Anthropic and OpenAI CEOs are pushing pacing plus independent evals for frontier systems.

Techniques

  • 1Try first

    Refuse answers sources do not support

    WORKS

    Stops shipping answers that look cited but are not backed.

    Agents often attach citations that do not actually support the claim. A grounding gate compares claim tokens to cited spans: accept when supported, refuse when citations are missing or unrelated, and regenerate when overlap is only partial. That blocks the common always-accept slip where a polished answer ships with decorative sources.

    ▸How it was tested

    Offline structural sim (no live model): 10 grounding cases. Gate matched all 10 accept/refuse/regenerate decisions (4 accept, 4 refuse, 2 regenerate); caught 6 cases a naive always-accept path would have slipped.

  • 2

    Do not swap tools on a hunch

    WORKS

    Blocks risky mid-flight tool swaps without proof.

    Mid-flight tool swaps are tempting when a call looks flaky, but suspicion alone is not enough. Require evidence that supports the failure hypothesis, a twin that passes structural checks (schema, reachability, side-effect class), and pairwise preference for the twin in both presentation orders so position bias cannot sneak a swap through.

    ▸How it was tested

    Offline structural sim (no live model): 8 replace cases. Gate matched all 8; allowed 2 full-gate replaces and blocked 6 (suspicion-only, missing signals, order bias, bad/unreachable/incomplete twin).

  • 3

    Check each tool step before next

    WORKS

    Stops bad tool steps from cascading into the rest of the run.

    A bad tool step that continues quietly poisons the rest of the trajectory. After each call, run four checks: does it advance the task, do used args match declared args, was execution valid, and were constraints satisfied. On fail, emit structured feedback and stop for replan instead of always-continuing into cascade.

    ▸How it was tested

    Offline structural sim (no live model): 8 trajectories. Stop-on-fail matched all 8 expected run lengths; blocked 6 cascades that naive always-continue completed. All four check dimensions fired on a deliberately broken step.

Sources

Links