Conversation
`run_agentic_what_if` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict. `--gate power` on a flaky item therefore reported a pass on the strength of one good run out of K -- the exact case the gate exists to catch, reporting the opposite of what happened. The evaluator now takes `gate` like every other multi-run kind, asks `gate_passed(...)`, and carries the gate's note into the assertion message, which matters here because the message body describes the BEST run and under pass^K that can be one that passed. It also stamps the gate on the run metadata and publishes pass@K/pass^K/gate_passed, so a gated run is readable in Langfuse rather than only in the exit code. The default is unchanged: without --gate this is pass@K exactly as before, pinned by a test alongside the power case. The power test fails against the previous code. Found by CodeRabbit on #1798, where the same gap was fixed for forecasting. `agentic_anomaly_detection` has it too and is fixed on its own branch, #1801, because that evaluator does not exist on master yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Warning Review limit reached
This review includes 3 billable files and costs up to $0.75. Or wait 33 minutes for your next included review. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (3)
Comment |
Tomkess
added a commit
that referenced
this pull request
Sep 24, 2026
`run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate power` on a flaky item reported a pass on the strength of one good run out of K. Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the assertion message -- which matters because the body describes the BEST run, and under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach Langfuse alongside the run metadata stamp. The default is unchanged: without --gate this is pass@K exactly as before, pinned by a test beside the power case. The power test fails against the previous code. Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is already on master, in #1831. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1831 +/- ##
=======================================
Coverage 83.08% 83.08%
=======================================
Files 330 330
Lines 21763 21767 +4
=======================================
+ Hits 18082 18086 +4
Misses 3681 3681 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Tomkess
added a commit
that referenced
this pull request
Sep 24, 2026
`run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate power` on a flaky item reported a pass on the strength of one good run out of K. Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the assertion message -- which matters because the body describes the BEST run, and under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach Langfuse alongside the run metadata stamp. The default is unchanged: without --gate this is pass@K exactly as before, pinned by a test beside the power case. The power test fails against the previous code. Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is already on master, in #1831. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> (cherry picked from commit 5632da8)
Tomkess
added a commit
that referenced
this pull request
Oct 2, 2026
…together Rebuilt from master rather than advanced: the branch's conversation.py predated master's multi-turn context work (QA-29448), while #1789 had already merged it, so merging master into the old tip would have meant hand-resolving a feature the PR branch already carried correctly. master + #1789 #1797 #1798 #1801 #1816 #1831 #1839. Three reconciliations the individual PRs cannot make on their own: - Dispatch registration in cli/agentic_runner.py is additive across four PRs that each add an evaluator; each pair conflicts and each resolution is the union. - #1816's structural test requires every multi-run evaluator to call build_failed_runs. dashboard_summary, forecasting and anomaly_detection postdate it and had no attachment point, so each grew one: a per-run detail function, build_failed_runs over the same predicate runs_passed is taken over, and failed_runs on both the outcome and the assertion error. dashboard_summary's _detail took the whole summary, so it is now a thin wrapper over a per-run _run_detail. - #1789 adds exit_reason/turns_used while #1816 moves the same dicts behind _run_detail. Both land: the per-run fields go into _run_detail, and max_iterations stays at the item level since it is the same for every run. Also supplies summary_input to #1816's failed-runs report test, which otherwise fails a dashboard-summary item on a missing fixture field before its evaluator is reached. 1520 passed, 1 skipped. ruff clean on everything these PRs touch; the two pre-existing format offenders under tests/ come from master untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess
added a commit
that referenced
this pull request
Oct 5, 2026
* feat(gooddata-eval): add the agentic anomaly-detection evaluator Scores the anomaly-detection skill end to end, with the granularity map aligned to what the service actually accepts and whole-conversation latency and cost rather than the goal turn's alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(gooddata-eval): let --gate decide an anomaly item, not pass@K alone `run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate power` on a flaky item reported a pass on the strength of one good run out of K. Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the assertion message -- which matters because the body describes the BEST run, and under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach Langfuse alongside the run metadata stamp. The default is unchanged: without --gate this is pass@K exactly as before, pinned by a test beside the power case. The power test fails against the previous code. Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is already on master, in #1831. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Tomkess
added a commit
that referenced
this pull request
Oct 5, 2026
Rebuilt on master now that #1789 (loop exit_reason) and #1801 (anomaly detection) have landed. Carries #1797, #1798, #1816, #1831, #1839. Three reconciliations the individual PRs cannot make on their own: - Dispatch registration in cli/agentic_runner.py is additive across the evaluator PRs; each pair conflicts and each resolution is the union. - #1789's exit_reason/turns_used are now on master in the same item-detail dicts #1816 moves behind a per-run builder. Both land: the per-run fields go into _run_detail, and max_iterations stays at the item level because it is the same for every run. - #1816's structural guard requires every multi-run evaluator to call build_failed_runs. dashboard_summary, forecasting and anomaly_detection postdate it and had no attachment point, so each grew one. anomaly detection is now ON MASTER without it, so this gap stops being a merge artefact the day #1816 lands. Also supplies summary_input to #1816's failed-runs report test, which otherwise fails a dashboard-summary item on a missing fixture field before its evaluator is reached. 1524 passed, 1 skipped. ruff clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess
added a commit
that referenced
this pull request
Oct 5, 2026
`run_agentic_dashboard_summary` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict, and dispatch never passed `gate` through at all. So `--gate power` on a flaky item returned a pass on the strength of one good run, and `run_agentic_items` then recorded gate_passed=True -- the report labels the run `power` while the item was decided by pass@K. Same defect already fixed in forecasting (#1798), anomaly detection (#1801) and what-if (#1831). A sweep of all twelve agentic evaluators says this was the last one carrying it. The gate also now reaches Langfuse: stamp_gate_metadata on the run and log_gate_scores per trace, so the published scores agree with the verdict instead of describing a gate the evaluator ignored. Found in review by CodeRabbit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess
added a commit
that referenced
this pull request
Oct 5, 2026
The same defect was found four times -- forecasting, anomaly detection, what-if and dashboard summary -- every time by a reviewer reading the diff, because nothing asserted the shape. An evaluator computes `pass_power_k` and then asks `if not summary.pass_at_k` for the verdict, so `--gate power` on a flaky item returns a pass while the report still labels the run `power`. Nothing looks broken; a flaky item has quietly been promoted. Structural, over every kind in AGENTIC_TEST_KINDS, skipping the ungated one. It lives here rather than in either fixing PR because neither carries both halves: #1797 fixes dashboard summary and #1831 fixes what-if, so the guard only passes where both are present. 1574 passed, 2 skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
run_agentic_what_ifcomputespass_power_k. The evaluator never read it:So
--gate poweron a flaky item passes on one good run out of K — the exact case the gate exists to catch, and the report says the opposite of what happened.This is live: what-if merged this morning in #1799.
The fix
What-if now does what the other eight multi-run kinds already do:
gate: EvalGate = DEFAULT_GATE, and the dispatch passes itgate_passed(...)for the verdictgate_failure_note(...)into the assertion message — this matters, because the message body describes the best run, which under pass^K can be a run that passed, so without the note a failing item reads like a passing oneDefault behaviour is unchanged. Without
--gatethis is pass@K exactly as before, pinned by its own test.Verification
Three tests: the power gate failing a 1-of-2 item,
anystill passing the same item, and the default matchingany. Reverting the one-line verdict change makes the power test fail, so it is testing the fix rather than the scaffolding.1209 tests pass,
ruff checkandruff format --checkclean.Related
Found by CodeRabbit on #1798, where the same gap is fixed for forecasting.
agentic_anomaly_detectionhas it too — fixed on #1801, since that evaluator is not on master yet. After those three, the only ungated kind isagentic_conversation, which is ungated by design.🤖 Generated with Claude Code