
TSK-1 · Anthropic · tested in Taskade Genesis
How Claude builds in Taskade Genesis
Use Claude to build apps in Taskade Genesis when you need polished apps that keep improving. Every TSK-1 Score comes from hands-on apps we opened, filled in, and asked to change. Open the scorecard, then compare models or start a build.
Updated · Last measured
TSK-1 scorecard for Claude
TSK-1 Score
Aggregated from hands-on Taskade Genesis builds · September 2026
- Interfaceready to share
- Strong
- Taskfollows your request
- Strong
- Memorykeeps your data
- Emerging
- Adapthandles follow-up edits
- Leading
Best for
Polished apps that keep improving
Claude produces the most polished finished apps we test and handles follow-up changes especially well. Its main risk is losing details in a long request.
Model results
Each Claude version below was tested by asking it to build an app in Taskade Genesis. Open one to read the findings, dated to the test that produced them.
Claude Sonnet 5Tested Sep 20267 findings
Our most carefully finished result in August. In September, one dashboard could not read its own sample data and a longer build stopped before the app was written.
Opened its own app and checked the assistant's answers — the only model to verify its work this way.
Most carefully finished result, with eight minor items left.
Built a contacts database instead of the requested sign-up form after two stalls.
Four attempts stalled before the sign-up form was completed.
Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.
Typed an 81-row scoring rubric and a 36-field sign-up table, then stopped at an approval prompt before writing the app.
Built a three-page match tracker whose dashboard reported zero matches over the eight matches it had entered itself.
Claude Opus 5Tested Sep 20266 findings
The most complete client sign-up build of September: five pages, every question in the customer's wording, and a scoring automation that ran and saved its verdict.
Set the design high-water mark with a polished, consistent interface.
Created the fullest workspace, with an assistant that understood its contents.
The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.
Best-executed match tracker of the test, without a single failed build action, at by far the highest cost.
Finished the client sign-up form with all 32 questions in the customer's wording, 31 word for word, five pages, and a scoring automation that saved its result.
Decided two open points in the brief and named both decisions in its closing summary.
Claude Haiku 4.5Tested Aug 20262 findings
Caught and repaired a problem that would have left the app blank.
Found the problem while building, repaired it, and then finished — the only model in the test to recover on its own.
Fastest build of the test, but it left out the weekly recap automation the brief asked for.
App kits you can open today
Live apps from the official Taskade account, not builds by Claude. Each is the same shape every benchmark request asks for: projects, agents, and automations you can open and clone.
Lead pipeline tracker CRMLead pipeline tracker4 projects · 2 agents · 3 flowsInvestor metrics dashboard DASHBOARDInvestor metrics dashboard5 projects · 2 agents · 4 flowsClient pipeline CRM CRMClient pipeline CRM2 projects · 1 agent · 2 flowsStore command center OPSStore command center34 projects · 1 agent · 2 flowsInvoice tracker TOOLInvoice tracker1 project · 1 agent · 2 flowsProduct launch dashboard DASHBOARDProduct launch dashboard1 project · 1 agent · 1 flow
Compare Claude
Other models in TSK-1
Same test, same four measures. Open any model to read its scorecard and findings.
FAQ
What is Claude best at in Taskade Genesis?
Claude is strongest when finish and follow-up changes matter. Sonnet produced our most carefully finished result, while Opus set the visual quality high-water mark. With long requests, review the finished app to make sure every detail carried through.
Claude Sonnet 5 or Opus 5 for app building?
Claude Opus 5 produced the more complete September sign-up app, with five pages and a scoring automation that ran. Claude Sonnet 5 remains strong at careful finish and follow-up changes, but review its dashboards against the saved data.
Why did Claude fail a test?
In one August test it built a contacts database instead of the client sign-up form, after two three-minute stalls lost it the thread. In another, four stalled runs in a row meant the client sign-up form never got built at all. We publish the tests that go badly alongside the ones that go well, and we open every app and read it back against the brief. That is the point of the benchmark.
Can I use Claude in Taskade?
Yes. Claude Sonnet 5, Opus 5, and Haiku 4.5 are all available in Taskade. Taskade offers 15+ frontier models from OpenAI, Anthropic, and open-weight providers.
Where can I compare Claude with other models, or try it myself?
Side-by-side model pages live at /compare. To try the same kind of app request yourself, start from /create. This page stays on what Claude produced in TSK-1 tests.




