Sitelet https://github.com/microsoft/AIOpsLab/pull/203
Skip to content

fix(orchestrator): keyword submit() calls are graded as an empty answer - #203

Open
Manuel Beardo Campo (Lostmanu) wants to merge 1 commit into
microsoft:mainfrom
Lostmanu:fix/submit-keyword-solution
Open

Manuel Beardo Campo (Lostmanu) wants to merge 1 commit into
microsoft:mainfrom
Lostmanu:fix/submit-keyword-solution

Conversation

@Lostmanu

Copy link
Copy Markdown

The bug

ResponseParser accepts submit(has_anomaly="Yes"), submit(faulty_components=[...]) and submit(analysis={...}) (covered by tests/parser/test_submit.py), and the submit actions accept them too, so the orchestrator replies VALID_SUBMISSION and ends the episode. But the solution is saved from positional arguments only (aiopslab/orchestrator/orchestrator.py:126-127):

if api == "submit":
    self.session.set_solution(args[0] if len(args) == 1 else args)

For a keyword call args is [], so the task is graded on an empty answer. With the real k8s_target_port_misconfig evals and the correct answer:

task positional call same call with the keyword, on main
detection Correct Invalid Format
localization 100.0, success 0.0, not successful
analysis success not successful

The same expression is in aiopslab/onboarding_evaluator.py:120, where the correct answer is rejected as INVALID_SUBMISSION instead.

The keyword form is not just theoretical. Agents see each submit's docstring with the parameter name (has_anomaly (str), faulty_components (list[str]), analysis (dict[str])), and at least one public agent built on AIOpsLab builds its detection and localization submissions in keyword form (GraphRCA, agents/archivist.py).

The fix

A small helper, submitted_solution() in aiopslab/utils/actions.py, binds the call to the task's submit() signature. A call that matches it stores the same value in either form. A call that does not match (unknown keyword, too many arguments, and so on) falls back to the previous expression, so it stores exactly what it stored before. Orchestrator.ask_env and the onboarding Evaluator.ask_env both use it.

Checked against main with the real actions and evals: the three keyword calls now grade like the positional ones, and five malformed calls (submit(answer="Yes"), submit("Yes", has_anomaly="No"), submit("Yes", "No"), submit(has_anomaly="Yes", extra=1), submit() on detection) store the same solution and get the same response as before.

Tests

tests/orchestrator/test_submit_solution.py runs ask_env of both runners with the real parsers and action classes, bypassing the cluster-dependent __init__ like test_mitigation_settle.py does. Its 6 keyword subtests fail on main and pass with this change. It also checks that a keyword submit() does not declare is still not taken as the answer, in both runners; this matters most in the onboarding evaluator, which grades the stored solution without calling the action. The unit tests were run on Python 3.12.14 with the project's dependencies and a dummy kubeconfig, since importing the package loads ~/.kube/config.

Side note: on main, 15 of the 21 tests in tests/parser/ fail. validate() from #84 expects the closing fence right after a newline, and those test inputs are indented triple-quoted strings. At 923b1be, just before #84 was merged, tests/parser/ passes. This PR does not touch them; happy to send a separate fix if that helps.

I do this because I'm passionate about it. I like helping, and finding bugs is a challenge I really enjoy.

The parsers accept submit(has_anomaly="Yes"), submit(faulty_components=[...])
and submit(analysis={...}) (tests/parser/test_submit.py covers them), but
Orchestrator.ask_env and the onboarding Evaluator only stored positional
arguments. A keyword submission was reported as valid and graded as empty:
detection "Invalid Format", localization 0.0, analysis not successful. The
onboarding evaluator rejected the correct answer instead.

Resolve the call through the task's submit() signature, so a call that
matches it stores the same solution in either form. Calls that do not match
the signature store exactly what they stored before.
@Lostmanu

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant