Clad Labs builds and trains production agents.

We do post-training and continual learning for production agents, so they get more accurate, faster and cheaper from their own traffic.

Get a free analysis

Applied research lab. YC F25. Founders from Caltech ML research and Meta.

Most agents ship once and stay the same. Every call after that is a record of what the agent saw, what it did and whether that was right. We turn that record into a model tuned to your work, and keep tuning it as traffic comes in.

How it works

01Traces

Your agent's production logs and the outcome signals that come with them: edits, corrections, accepts, rejects.

02Grader

A rubric and a judge, calibrated against your labels, so a miss means what your team says it means.

03Train

Post-training on your own traffic. The candidate model runs against the current one on held-out tasks before anything changes.

04Deploy and repeat

The new model goes behind your endpoint. New traffic feeds the next cycle.

You own the model, the evals and the endpoint.

For companies running agents in production: coding, voice, support, document and vertical workflows. One step of one agent is enough to start.

Free analysis

Send us about 200 recent traces from one step of your agent. In a week you get the failure rate, the main failure modes, and a roadmap to fix them. Zero work on your side beyond the export.

Get a free analysis