You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Keep the Basecamp skill focused on durable operating and safety rules instead of duplicating the CLI's command manual.
The existing skill had grown to 89 KB / 1,565 lines. Loading it for a simple question such as “how many Basecamp accounts do I have?” added roughly 13,000 tokens of tool output before the command could run, even though the installed CLI already exposes structured, command-specific --agent --help.
What changed
reduce the main skill to 8.3 KB / 235 lines (about 91% smaller)
make live basecamp ... --agent --help the source of truth for command syntax, flags, scope, and notes
retain cross-cutting rules for credentials, URLs and replies, output modes, project/account scope, stdin content, attachments, retries, authentication, and mutations
teach the eval harness to serve structured help from the locally built CLI
add an eval for the original account-count question
enforce a 16 KiB skill budget and discovery-first contract in Go tests
document deterministic checks and model-eval comparison workflows
Why the eval harness changed
A discovery-first skill cannot be evaluated faithfully if help calls receive the generic empty mock. The harness now executes only structured --agent --help requests against the local binary; normal Basecamp operations remain deterministic mocks.
Verification
GOWORK=off bin/ci
eval patterns: 162 compiled across 30 cases
structured-help harness self-test
skill drift checks
The hosted model eval was not run locally because ANTHROPIC_API_KEY is not configured; the existing PR eval workflow can run it when the repository secret is available.
Summary by cubic
Slims the Basecamp skill from 89 KB / 1,565 lines to 8.3 KB / 235 lines by removing the embedded command manual and treating live basecamp ... --agent --help output as the source of truth for syntax, flags, and scope. Loading the old skill added roughly 13,000 tokens for simple questions; the new one keeps only durable operating and safety rules like credentials, output modes, and mutation retries. Discovery now goes leaf-first: accounts list --agent --help directly instead of walking down the command tree.
Eval harness
Serves structured --agent --help from the locally built CLI and enforces leaf-first discovery; other Basecamp operations remain deterministic mocks.
Adds eval cases for the account-count question and Docs & Files folder downloads, with tightened mock and matcher regexes that enforce correct command and flag sequences, plus a self-test for the structured help path, runnable via make check-eval-harness.
Go tests enforce a 16 KiB skill budget and the discovery-first contract.
Clarifies that the todolist --description flag accepts rich text HTML, with a test guarding the updated help text.
Written for commit 575a735. Summary will update on new commits.
Slims the Basecamp skill by replacing its command catalog with live structured CLI help, while retaining durable safety guidance.
Changes:
Introduces a discovery-first skill with a 16 KiB budget.
Serves live structured help in evals and adds an account-count case.
Documents and integrates deterministic harness checks.
[!TIP]
If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or run gh pr ready --undo.
Click "Ready for review" or run gh pr ready to reengage.
File
Description
skills/basecamp/SKILL.md
Replaces the embedded manual with durable operating rules.
Fixed — structured help now routes only allowlisted discovery requests through the non-mutating help command, and retry guidance now distinguishes an absent retry signal from a definitive verdict.
Run deterministic eval harness self-test in integration workflow
Makefile:415
The remote integration workflow invokes its checks individually and currently runs make check-eval-patterns, but neither make check nor make check-eval-harness. Consequently, whenever the secret-gated model-eval job is skipped, CI never runs this new deterministic harness self-test. Add make check-eval-harness to the integration workflow so the advertised no-credential check is enforced for every PR.
Accept valid direct-count and filtered-list answers in evaluation
skill-evals/cases/accounts-count.yml:8
This expectation rejects valid answers that live root help can lead the model to: accounts list --count directly emits the list length, and accounts list --agent --jq 'length' correctly filters the data-only payload. Either invocation answers the task after introspection, but the eval would fail despite receiving and returning 2. Accept those forms as well as envelope-based .data | length.
The reason will be displayed to describe this comment to others. Learn more.
1 issue found across 2 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="skill-evals/cases/accounts-count.yml">
<violation number="1" location="skill-evals/cases/accounts-count.yml:14">
P3: `max_commands: 2` leaves zero headroom for the skill's own fallback path from the new workflow. Step 3 of the rewritten SKILL.md tells agents that a leaf-miss (a plausible first guess here, e.g. `accounts count --agent --help`) is recovered by inspecting the parent group, which makes a fully skill-conformant run take 3 commands: wrong leaf help, parent help, then the count. With one sample per case, any first-guess miss or error-reading step flunks the case even though the behavior matches the changed skill exactly.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
The reason will be displayed to describe this comment to others. Learn more.
P3: max_commands: 2 leaves zero headroom for the skill's own fallback path from the new workflow. Step 3 of the rewritten SKILL.md tells agents that a leaf-miss (a plausible first guess here, e.g. accounts count --agent --help) is recovered by inspecting the parent group, which makes a fully skill-conformant run take 3 commands: wrong leaf help, parent help, then the count. With one sample per case, any first-guess miss or error-reading step flunks the case even though the behavior matches the changed skill exactly.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At skill-evals/cases/accounts-count.yml, line 14:
<comment>`max_commands: 2` leaves zero headroom for the skill's own fallback path from the new workflow. Step 3 of the rewritten SKILL.md tells agents that a leaf-miss (a plausible first guess here, e.g. `accounts count --agent --help`) is recovered by inspecting the parent group, which makes a fully skill-conformant run take 3 commands: wrong leaf help, parent help, then the count. With one sample per case, any first-guess miss or error-reading step flunks the case even though the behavior matches the changed skill exactly.</comment>
<file context>
@@ -4,8 +4,11 @@ mocks:
accept_response:
- '\b2\b'
-max_commands: 4
+max_commands: 2
</file context>
CI skips the deterministic structured-help self-test
Makefile:415
This target is only added to the local aggregate check, but GitHub Actions does not invoke make check: .github/workflows/test.yml runs make check-eval-patterns directly, and the model-eval job is skipped when ANTHROPIC_API_KEY is unavailable. As a result, the new deterministic structured-help self-test can regress while required CI stays green. Add make check-eval-harness to the deterministic workflow job as well.
Flagged for human review — wiring a new required GitHub Actions step changes security-sensitive CI configuration, which this review-processing workflow does not modify.
Not doing this — the two-command ceiling is the deliberate efficiency contract for this obvious account-list case; increasing it would stop the eval from detecting the discovery overhead it was created to catch.
Fixed — the account-count eval now validates the exact counting expression for envelope and data-only modes instead of accepting any jq filter containing length.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
commandsCLI command implementationsskillsAgent skillstestsTests (unit and e2e)
2 participants
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Keep the Basecamp skill focused on durable operating and safety rules instead of duplicating the CLI's command manual.
The existing skill had grown to 89 KB / 1,565 lines. Loading it for a simple question such as “how many Basecamp accounts do I have?” added roughly 13,000 tokens of tool output before the command could run, even though the installed CLI already exposes structured, command-specific
--agent --help.What changed
basecamp ... --agent --helpthe source of truth for command syntax, flags, scope, and notesWhy the eval harness changed
A discovery-first skill cannot be evaluated faithfully if help calls receive the generic empty mock. The harness now executes only structured
--agent --helprequests against the local binary; normal Basecamp operations remain deterministic mocks.Verification
GOWORK=off bin/ciThe hosted model eval was not run locally because
ANTHROPIC_API_KEYis not configured; the existing PR eval workflow can run it when the repository secret is available.Summary by cubic
Slims the Basecamp skill from 89 KB / 1,565 lines to 8.3 KB / 235 lines by removing the embedded command manual and treating live
basecamp ... --agent --helpoutput as the source of truth for syntax, flags, and scope. Loading the old skill added roughly 13,000 tokens for simple questions; the new one keeps only durable operating and safety rules like credentials, output modes, and mutation retries. Discovery now goes leaf-first:accounts list --agent --helpdirectly instead of walking down the command tree.Eval harness
--agent --helpfrom the locally built CLI and enforces leaf-first discovery; other Basecamp operations remain deterministic mocks.make check-eval-harness.--descriptionflag accepts rich text HTML, with a test guarding the updated help text.Written for commit 575a735. Summary will update on new commits.