fix(orchestrator): keyword submit() calls are graded as an empty answer - #203
Open
Manuel Beardo Campo (Lostmanu) wants to merge 1 commit into
Open
Manuel Beardo Campo (Lostmanu) wants to merge 1 commit into
Manuel Beardo Campo (Lostmanu) wants to merge 1 commit into
Conversation
The parsers accept submit(has_anomaly="Yes"), submit(faulty_components=[...])
and submit(analysis={...}) (tests/parser/test_submit.py covers them), but
Orchestrator.ask_env and the onboarding Evaluator only stored positional
arguments. A keyword submission was reported as valid and graded as empty:
detection "Invalid Format", localization 0.0, analysis not successful. The
onboarding evaluator rejected the correct answer instead.
Resolve the call through the task's submit() signature, so a call that
matches it stores the same solution in either form. Calls that do not match
the signature store exactly what they stored before.
Author
|
@microsoft-github-policy-service agree |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
ResponseParseracceptssubmit(has_anomaly="Yes"),submit(faulty_components=[...])andsubmit(analysis={...})(covered bytests/parser/test_submit.py), and the submit actions accept them too, so the orchestrator repliesVALID_SUBMISSIONand ends the episode. But the solution is saved from positional arguments only (aiopslab/orchestrator/orchestrator.py:126-127):For a keyword call
argsis[], so the task is graded on an empty answer. With the realk8s_target_port_misconfigevals and the correct answer:The same expression is in
aiopslab/onboarding_evaluator.py:120, where the correct answer is rejected asINVALID_SUBMISSIONinstead.The keyword form is not just theoretical. Agents see each submit's docstring with the parameter name (
has_anomaly (str),faulty_components (list[str]),analysis (dict[str])), and at least one public agent built on AIOpsLab builds its detection and localization submissions in keyword form (GraphRCA, agents/archivist.py).The fix
A small helper,
submitted_solution()inaiopslab/utils/actions.py, binds the call to the task'ssubmit()signature. A call that matches it stores the same value in either form. A call that does not match (unknown keyword, too many arguments, and so on) falls back to the previous expression, so it stores exactly what it stored before.Orchestrator.ask_envand the onboardingEvaluator.ask_envboth use it.Checked against main with the real actions and evals: the three keyword calls now grade like the positional ones, and five malformed calls (
submit(answer="Yes"),submit("Yes", has_anomaly="No"),submit("Yes", "No"),submit(has_anomaly="Yes", extra=1),submit()on detection) store the same solution and get the same response as before.Tests
tests/orchestrator/test_submit_solution.pyrunsask_envof both runners with the real parsers and action classes, bypassing the cluster-dependent__init__liketest_mitigation_settle.pydoes. Its 6 keyword subtests fail on main and pass with this change. It also checks that a keywordsubmit()does not declare is still not taken as the answer, in both runners; this matters most in the onboarding evaluator, which grades the stored solution without calling the action. The unit tests were run on Python 3.12.14 with the project's dependencies and a dummy kubeconfig, since importing the package loads~/.kube/config.Side note: on main, 15 of the 21 tests in
tests/parser/fail.validate()from #84 expects the closing fence right after a newline, and those test inputs are indented triple-quoted strings. At 923b1be, just before #84 was merged,tests/parser/passes. This PR does not touch them; happy to send a separate fix if that helps.I do this because I'm passionate about it. I like helping, and finding bugs is a challenge I really enjoy.