Sitelet https://github.com/martinfrancois/java-streams-skill/pull/94
Skip to content

fix(skill): tighten the stream workflow for the current Tessl reviewer - #94

Merged
martinfrancois merged 3 commits into
mainfrom
fix/skill-review-100
Oct 4, 2026
Merged

martinfrancois merged 3 commits into
mainfrom
fix/skill-review-100

Conversation

@martinfrancois

@martinfrancois martinfrancois commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Problem: the current Tessl reviewer scores the released skills/java-streams/SKILL.md at 91 (content 79), so the Tessl skill review check is red on every pull request (fix(skill): released SKILL.md scores 91% under the current Tessl reviewer #88).
  • Why it matters: the workflow gate is 100, and the skill review is part of what the registry shows.
  • What changed: the opening paragraph is four lines instead of six; the find-first rule lives in the decision table plus a checkpoint after API selection; step 2 and step 4 each carry a one-line snippet; a compact toMap example with merge, null filtering, and an identity mapper joins step 4; the review-output bullets are one imperative sentence. Every rule keeps its meaning.
  • What did not change: the references and the evals.
  • October follow-up (commit a81acce): the reviewer changed again and scored the first rewrite 94. The second commit states the named-helper rule once (step 5), shows Gatherers.mapConcurrent as a complete pipeline, wraps the toMap merge example in the SeatIndex class it references, drops the OrderChecks snippet that repeated the anyMatch advice, points to the parallel-safety conditions once (step 7), and scopes the generic trigger words in the description to stream stages. No rule was added or removed.

Quality review on this branch: 100 (01a10390-fe22-72b8-bfe2-26ecb066c104), after 94, 97, 97, and 98 on the October drafts and one N/A when a colon in the description broke the YAML front matter.

Change Type

  • Skill behavior
  • Evals or scoring
  • Documentation
  • CI, release, or dependency automation
  • Repository metadata or contribution process
  • Other maintenance

Linked Issue

User-Visible Behavior

Same guidance, shorter body, one more copy-paste example.

Bug Fix Details

  • Root cause: Tessl replaced the single-pass skill review with a full-bundle agent review that weighs conciseness and actionability differently.
  • Test, eval, or guardrail added: none; the existing suites are the guardrail and must rerun (below).
  • If no test or eval was added, why not: the change is wording and examples.

Ownership Boundary

Unchanged. The example uses Function.identity() inside a collector, which the ownership page lists as stream-owned collector semantics.

Hosted Evidence

Lift-sensitive: this changes the measured skill bundle. Revert strategy: revert the squash commit. Every run below used the runtime text of commit a81acce and the Tessl default solver, which changed from deepseek-v4-flash to deepseek-v4.1-flash before this window.

Check Result
tessl review run --threshold 100 skills/java-streams/SKILL.md 100, run 01a10390-fe22-72b8-bfe2-26ecb066c104
Probes main 01, 02 (both variants) with 100, without 100, run 01a10394-27cb-711f-8632-12d5b0c0b1e9
Probes regression 04, 22, 25 (with context) all 100, run 01a1039a-d4e9-76d2-9e7a-56945dfa7352
Main 03, 04 (both variants) with 100, without 100, run 01a103a1-b700-720e-abaa-84f17fbb23c2
Regression, the other 16 (with context) all 100, run 01a103a9-22fe-713b-bec2-f1bf09957918
Reference, all 6 (with context only, to save credits) all 100, run 01a103b0-1209-7449-b98b-7c573a67ece7
Composition with java-functional-style PR #18 text, main 01, 02, 03 100; 04 130/200 (two stream passes instead of teeing) in 01a10394-02b3-72dc-beb0-cea488f2597f; isolated rerun of 04 200/200 in 01a103ae-a799-73ca-b347-d515c7fa845d

Every scenario reached 100 with context. The 2.22x bar from the handoff cannot be met: all four main baselines also scored 100 under the new solver, so the main suite shows 1.0x. That comes from the solver, not the skill text, and it needs a maintainer decision (new main scenarios, or accept that the registry score drops) before release. Results are also recorded in the NUMBERING.md notes. The per-scenario criteria-meta.json sidecars exist only on main, so they get updated after this branch is rebased.

Before merge: the ruleset requires an up-to-date branch, and this branch is behind main. Rebase it onto main (no conflicts; skills/ is unchanged on main since the base) and force-push with lease.

Release note: v1.2.1 is a GitHub-only release; its Tessl publish failed at the review gate (run 35561430400), so the registry stays at 1.2.0 until this PR lands as 1.2.2.

Validation

Checks most contributors can run:

  • python3 scripts/validate_skill.py skills/java-streams -> passed
  • python3 scripts/validate_eval_criteria.py evals evals-reference evals-regression -> passed, 29 scenarios
  • python3 -m py_compile scripts/*.py -> passed
  • bash -n scripts/*.sh -> passed
  • python3 scripts/validate_json_files.py -> passed
  • python3 scripts/validate_openai_agent_yaml.py -> passed
  • tessl plugin lint . -> passed
  • Manual rendered-doc review

Tessl-authenticated checks:

  • tessl review run --workspace martinfrancois --threshold 100 skills/java-streams/SKILL.md -> 100
  • Hosted evals -> see the Hosted Evidence table

Human Verification

Each rewritten sentence was checked against the original for a dropped rule: none dropped; the find-first exception, the min/max caveat, the parallelStream quote, and the Java 24 gatherer shape are all still stated.

Review Checklist

  • The change is scoped to the sections, skill files, evals, or workflows described above.
  • Validation that applies to this change is checked above, or any unavailable check is explained.
  • If Java stream guidance changed, Java baseline compatibility plus ordering, null handling, and parallelism were considered.
  • If evals or benchmark claims changed, the eval scenarios remain fair and do not leak answer keys, run IDs, or fixed score claims into runtime references. (No eval change.)
  • If runtime skill text or references changed, hosted checks were widened as described in docs/agents/workflow.md, or any Tessl blocker is documented.
  • If a runtime skill/reference change was released, the final report includes the published main eval run plus post-change reference and regression run IDs, or a blocker issue for missing broad suites. (Not released.)
  • Main and reference evals were run with both variants when hosted evals were needed; regression evals were run with context only unless reclassification back to reference was being checked. (Main both variants; reference with context only, to save credits.)
  • New or moved eval scenarios follow the classifier recommendation, or the PR explains the maintainer-approved override. (None.)
  • Every retained eval scenario has a 100% with-context result, or any below-100 result is documented as blocking follow-up rather than classified/reportable coverage.
  • PR title or squash title uses Conventional Commits.
  • Redaction checked: no tokens, private links, private eval artifacts, local host paths, or proprietary Java source.

AI Assistance (if used)

  • AI-assisted PR
  • I confirm I understand and reviewed the change (the October follow-up commits are not yet reviewed by the maintainer)

@martinfrancois

Copy link
Copy Markdown
Owner Author

Handoff for the hosted proof and the 1.2.2 release: #95 (do on or after 1 October 2026, after java-functional-style-skill 1.0.0 and the optionals release).

Closes #88 once the hosted reruns pass. The current Tessl reviewer
scores the released SKILL.md at 91 (content 79): the opening paragraph
and the find-first prose were dense, the second code block was a
one-line fragment, and verification sat only in the last step. The
opening is shorter, the find-first rule lives in the decision table
with a checkpoint after API selection, and a complete toMap example
with merge, null filtering, and an identity mapper joins the workflow.

Lift-sensitive: this changes the measured skill bundle; every suite
and the composition check must rerun before release.
The reviewer changed after September and scored the previous rewrite 94:
it flagged the named-helper rule repeated across steps 3 to 5, the
Gatherers.mapConcurrent line built from placeholder calls, and a toMap
example that referenced SeatIndex::later without a SeatIndex class.

- state the named-helper rule once, in step 5
- show mapConcurrent as a complete pipeline that carries each element
  with its result
- wrap the merge example in the SeatIndex class it references
- drop the OrderChecks snippet, which repeated the anyMatch guidance
- point to the parallel-safety conditions once, from step 7
- scope the generic trigger terms in the description to stream stages

No rule was added or removed. Review run 01a10390-fe22-72b8-bfe2-26ecb066c104
scores 100.
Main with both variants, reference and regression with context, and the
composition check with java-functional-style ran against commit a81acce.
Every scenario reached 100 with context. The main baselines also reached
100 under the new Tessl default solver, which the main notes record.
@martinfrancois
martinfrancois merged commit eaea5a1 into main Oct 4, 2026
8 checks passed
@martinfrancois
martinfrancois deleted the fix/skill-review-100 branch October 4, 2026 01:52
martinfrancois pushed a commit that referenced this pull request Oct 4, 2026
🤖 I have created a release *beep* *boop*
---


##
[1.2.2](v1.2.1...v1.2.2)
(2026-10-04)


### Bug Fixes

* **skill:** tighten the stream workflow for the current Tessl reviewer
([#94](#94))
([eaea5a1](eaea5a1))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
martinfrancois added a commit that referenced this pull request Oct 4, 2026
…110)

## Summary

- Problem: Renovate automerged the Tessl CLI bumps in #97, #103 and #104
while `Tessl skill review` was failing. The `Required merge checks`
ruleset only requires `Validate skill and plugin` and `Commitlint`.
- Why it matters: a Tessl update that lowers the review score below the
100 threshold lands on main without anyone seeing it. #98 tried to undo
one of these by pinning the CLI back; it is closed now that #94 brought
the score back to 100.
- What changed: `Tessl skill review` now reports on every PR, so the
ruleset can require it. `skill-review.yml` has no paths filter anymore;
a step checks whether the PR changes a file the review reads (the same
list the filter had) and skips the review otherwise, so the job still
passes. Release Please PRs do not trigger `pull_request` workflows, so
`release-please.yml` runs the review on the release PR and posts the
`Tessl skill review` status, as it already does for the other two
required checks. `docs/agents/workflow.md` lists the new status.
- What did not change: the review command, its threshold, the skill, and
the evals. The ruleset itself is not part of this PR; I add `Tessl skill
review` to it after this merges, so PRs opened before the change do not
get stuck.

## Change Type

- [ ] Skill behavior
- [ ] Evals or scoring
- [x] Documentation
- [x] CI, release, or dependency automation
- [ ] Repository metadata or contribution process
- [ ] Other maintenance

## Linked Issue

- Related #97, #98, #103, #104

## User-Visible Behavior

After the ruleset update, a PR with a failing Tessl skill review cannot
merge, including Renovate automerges. PRs that touch none of the
reviewed files show the check as passed with a skip message.

## Bug Fix Details

- Root cause: the skill review ran but was not a required check, and it
could not be made required as it was, because its paths filter and the
Release Please token meant many PRs would never get the check.
- Test, eval, or guardrail added: the required check itself.
- If no test or eval was added, why not: workflow logic; verified with
actionlint and by running the detection step against synthetic commits
(below).

## Validation

Checks most contributors can run:

- [x] `python3 scripts/validate_skill.py skills/java-streams`
- [x] `python3 scripts/validate_eval_criteria.py evals evals-reference
evals-regression`
- [x] `python3 -m py_compile scripts/*.py`
- [x] `bash -n scripts/*.sh`
- [x] `tessl plugin lint .`
- [ ] Manual rendered-doc or example review, if docs or examples changed

Tessl-authenticated checks:

- [x] `bash scripts/check_publish_dry_run.sh .`
- [x] `tessl plugin publish --dry-run --bump patch .`
- [ ] `tessl review run ...`: skill text unchanged; this PR's own `Tessl
skill review` run covers it because it touches `skill-review.yml`.
- [ ] Hosted evals: not needed, no skill or eval change.

Details:

```text
scripts/validate_repo.sh: passed (all of the first five checks above)
actionlint 1.7.12 with shellcheck 0.11.0 on .github/workflows/*.yml: no findings
Detection step run against synthetic merge commits:
  package.json                  -> changed=false
  .github/workflows/ci.yml      -> changed=false
  skills/java-streams/SKILL.md  -> changed=true
  docs/agents/evals.md          -> changed=true
publish dry-runs: passed (would bump 1.2.2 to 1.2.3)
```

## Human Verification

```text
The release PR path can only run on the next Release Please PR. Check that it shows a
"Tessl skill review" status next to "Validate skill and plugin" and "Commitlint".
```

## Review Checklist

- [x] The change is scoped to the sections, skill files, evals, or
workflows described above.
- [x] Validation that applies to this change is checked above, or any
unavailable check is explained.
- [ ] If Java stream guidance changed, Java baseline compatibility plus
ordering, null handling, and parallelism were considered.
- [ ] If evals or benchmark claims changed, the eval scenarios remain
fair and do not leak answer keys, run IDs, or fixed score claims into
runtime references.
- [ ] If runtime skill text or references changed, hosted checks were
widened from targeted affected scenarios to main/reference/regression as
described in `docs/agents/workflow.md`, or any Tessl blocker is
documented.
- [ ] If a runtime skill/reference change was released, the final report
includes the published main eval run plus post-change reference and
regression run IDs, or a blocker issue for missing broad suites.
- [ ] Main and reference evals were run with both variants when hosted
evals were needed; regression evals were run with context only unless
reclassification back to reference was being checked.
- [ ] New or moved eval scenarios follow the classifier recommendation,
or the PR explains the maintainer-approved override.
- [ ] Every retained eval scenario has a 100% with-context result, or
any below-100 result is documented as blocking follow-up rather than
classified/reportable coverage.
- [x] PR title or squash title uses Conventional Commits.
- [x] Redaction checked: no tokens, private links, private eval
artifacts, local host paths, or proprietary Java source.

## AI Assistance (if used)

- [x] AI-assisted PR
- [ ] I confirm I understand and reviewed the change

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(skill): released SKILL.md scores 91% under the current Tessl reviewer

1 participant