download dots

TSK-1 · Anthropic · tested in Taskade Genesis

How Claude builds in Taskade Genesis

Use Claude to build apps in Taskade Genesis when you need polished apps that keep improving. Every TSK-1 Score comes from hands-on apps we opened, filled in, and asked to change. Open the scorecard, then compare models or start a build.

Updated · Last measured

Your requestTaskade EVETSK-1Workspace memoryLive app

TSK-1 scorecard for Claude

TSK-1 Score

Aggregated from hands-on Taskade Genesis builds · September 2026

TSK-1 Score 75 out of 100
Interfaceready to share
Strong
Taskfollows your request
Strong
Memorykeeps your data
Emerging
Adapthandles follow-up edits
Leading

Best for

Polished apps that keep improving

Claude produces the most polished finished apps we test and handles follow-up changes especially well. Its main risk is losing details in a long request.

Model results

Each Claude version below was tested by asking it to build an app in Taskade Genesis. Open one to read the findings, dated to the test that produced them.

Claude Sonnet 5Tested Sep 20267 findings

Our most carefully finished result in August. In September, one dashboard could not read its own sample data and a longer build stopped before the app was written.

Jul 30, 2026

Opened its own app and checked the assistant's answers — the only model to verify its work this way.

Aug 3, 2026

Most carefully finished result, with eight minor items left.

Aug 3, 2026

Built a contacts database instead of the requested sign-up form after two stalls.

Aug 6, 2026

Four attempts stalled before the sign-up form was completed.

Aug 25, 2026

Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.

Sep 19, 2026

Typed an 81-row scoring rubric and a 36-field sign-up table, then stopped at an approval prompt before writing the app.

Sep 19, 2026

Built a three-page match tracker whose dashboard reported zero matches over the eight matches it had entered itself.

Claude Opus 5Tested Sep 20266 findings

The most complete client sign-up build of September: five pages, every question in the customer's wording, and a scoring automation that ran and saved its verdict.

Jul 30, 2026

Set the design high-water mark with a polished, consistent interface.

Aug 1, 2026

Created the fullest workspace, with an assistant that understood its contents.

Aug 25, 2026

The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.

Aug 25, 2026

Best-executed match tracker of the test, without a single failed build action, at by far the highest cost.

Sep 19, 2026

Finished the client sign-up form with all 32 questions in the customer's wording, 31 word for word, five pages, and a scoring automation that saved its result.

Sep 19, 2026

Decided two open points in the brief and named both decisions in its closing summary.

Claude Haiku 4.5Tested Aug 20262 findings

Caught and repaired a problem that would have left the app blank.

Jul 31, 2026

Found the problem while building, repaired it, and then finished — the only model in the test to recover on its own.

Aug 25, 2026

Fastest build of the test, but it left out the weekly recap automation the brief asked for.

App kits you can open today

Live apps from the official Taskade account, not builds by Claude. Each is the same shape every benchmark request asks for: projects, agents, and automations you can open and clone.

Browse all App Kits →

Compare Claude

Other models in TSK-1

Same test, same four measures. Open any model to read its scorecard and findings.

See the full TSK-1 ranking →

FAQ

What is Claude best at in Taskade Genesis?

Claude is strongest when finish and follow-up changes matter. Sonnet produced our most carefully finished result, while Opus set the visual quality high-water mark. With long requests, review the finished app to make sure every detail carried through.

Claude Sonnet 5 or Opus 5 for app building?

Claude Opus 5 produced the more complete September sign-up app, with five pages and a scoring automation that ran. Claude Sonnet 5 remains strong at careful finish and follow-up changes, but review its dashboards against the saved data.

Why did Claude fail a test?

In one August test it built a contacts database instead of the client sign-up form, after two three-minute stalls lost it the thread. In another, four stalled runs in a row meant the client sign-up form never got built at all. We publish the tests that go badly alongside the ones that go well, and we open every app and read it back against the brief. That is the point of the benchmark.

Can I use Claude in Taskade?

Yes. Claude Sonnet 5, Opus 5, and Haiku 4.5 are all available in Taskade. Taskade offers 15+ frontier models from OpenAI, Anthropic, and open-weight providers.

Where can I compare Claude with other models, or try it myself?

Side-by-side model pages live at /compare. To try the same kind of app request yourself, start from /create. This page stays on what Claude produced in TSK-1 tests.