Why the trajectory and not a threshold? A single score compared against a fixed floor is a weak trigger for an agent mid-task: early in any real task the work genuinely is incomplete, so a low score is correct rather than a fault. Judging the series instead — is it climbing, stalled, or falling back? — separates an agent making progress from one that is stuck, which a point comparison cannot do.
The three tiers
Evaluation is a ladder. Cheap, always-on signals watch progress; expensive, surgical signals run only once the cheap ones say something is wrong.
Tier 1 is deliberately cheap:
task_completion is the outcome trajectory and coherence is an embedding-based (free) co-gate. Tier 3 is opt-in because every LLM-judge metric reads the whole trace, so cost scales with metric count.
These are the metrics the platform reports as runnable via
pandaprobe evals metrics --target trace. The set is configurable per tier if your project defines its own metrics.The trajectory gate
A Tier-1 metric does not breach on its absolute value. An agent three steps into a task has legitimately not completed it, so a low earlytask_completion is expected, not a fault. What matters is the shape of the series:
STALL — no progress toward the target
STALL — no progress toward the target
gate_window consecutive traces (default 5) pass without a new high, and the session’s best score so far is still below gate_target (default 0.5). The agent is working but not getting closer to done.Note it is the peak that is compared against gate_target, not the latest score: once a session has genuinely reached the target, a later flat stretch is not a stall.REGRESSION — a fall from the running peak
REGRESSION — a fall from the running peak
The score drops more than
gate_drop (default 0.15) below the best value this session has reached. The agent had it and lost it. This one applies regardless of gate_target.RESET ON GAIN — the healthy case
RESET ON GAIN — the healthy case
A score reaching
peak + gate_gain (default +0.02) raises the peak and resets the window to zero. A session that keeps improving never breaches, no matter how low it starts or how long it takes.Progress is measured against the peak, not the previous trace — so oscillating around a plateau does not read as progress.state/score_history.json, so it survives process restarts, and a whole trace’s metrics fold into it in a single write.
Which tiers can breach on a value
Only tiers 0 and 2 treat “below threshold” as a finding:Severity: advisory vs. actionable
The mapping is deliberate, because it decides what becomes a promotable training signal:
Only a confirmed, surgically-diagnosed failure becomes a replayable eval case. A trajectory wobble the step-level metrics can’t corroborate is worth telling the agent about, but not worth training on.
Conditions and signatures
Every condition a score satisfies becomes a label, and a signature iscondition:metric:
So
stall:task_completion and breach:tool_correctness are signatures. They drive de-duplication, rule tagging, eval-case matching, and forward-trial baselines throughout the harness.
How a turn is scored
- One platform run per turn for Tier 1, however many traces landed — the batch endpoint takes every id at once and tags each score with its trace. The gate’s fold is still applied one trace at a time in chronological order, because a peak/stall fold is history-dependent; only the evals are batched.
- Polling is bounded by
poll_interval_s×poll_max_attempts. - Trace ingestion lags turn-end (the SDK flushes on a background thread), so transiently empty runs are retried with backoff (
eval_retry_attempts,eval_retry_backoff_s). An empty trace listing is only retried while the session has never yet produced a trace — once one has landed, an empty listing is the truth, and retrying it would sleep out the backoff budget on every healthy turn with nothing new to score. - Any persistent CLI failure degrades to a pending score. A pending score is never a breach, and the harness never raises into, or blocks, your agent loop.
Before the first evaluation, a memoized health check verifies the CLI is present and authenticated (
pandaprobe version + auth status). On failure the harness runs degraded: one warning, a journal health event, evaluations skipped — never a crash, never a silent no-op.The per-turn barrier
Detection is worthless if the lesson arrives after the task is over.harness.settle(session_id) blocks until this turn’s diagnosis has landed — the evaluation resolved, the report handled, the eval case captured, the notice posted — so the agent’s next turn sees the mailbox and rule set the harness just produced:
barrier_timeout_s (default 180s), deliberately separate from drain_timeout_s, which is only a best-effort join. On expiry the work stays running detached and timed_out is set: a slow platform degrades the loop’s latency, never its correctness.
hook.pending_sessions exposes which sessions still have work in flight, for host-side phase barriers (“has everything landed before I archive the workspace?”).
De-duplication, cooldown, and recovery
A persistent problem should post one notice, not one per turn:- Per session, the harness remembers the last posted signature set. A new notice posts only when a new condition appears — or, with
alert_cooldown_turns > 0, when the cooldown expires while the same conditions persist. - When a previously-alerting session scores clean, the state resets and a
recoveryevent is journaled. A later regression of the same kind alerts again.
The circuit breaker
If something goes systemically wrong — a bad deploy, a broken tool — notices can storm. More thancircuit_breaker_max_notices within circuit_breaker_window_s escalates to a single needs_human notice and suppresses further posting until the window drains. The standing protocol instructs the agent to surface that notice to a human rather than act on it: it is the one deliberate exit from the autonomous loop.
Shadow mode
observe_only=true evaluates and journals everything but never posts to the mailbox — useful for tuning the gate against real traffic before letting the agent act. Pair it with the calibration CLI.
