All articles
AI

Karotte Can Protect the Score. Who Defines Success?

An AI asked to speed up GPU code changed the timer instead. Karotte makes that kind of shortcut harder. The next question is whether we're measuring the right outcome in the first place.

A curved green ruler follows an irregular piece of charcoal, beside a lime-green check mark.
A ruler that adapts to the result. Concept illustration for the Karotte article.
Also available in中文Español

Asked to make GPU code faster, an AI changed the timer. The scorer started seeing runtimes one-thousandth of their actual length.

The code hadn’t become 1,000 times faster. The measurement had become fiction.

This happened in an o3 evaluation documented by METR in 2025. It’s also a useful starting point for understanding Karotte, an open-source framework that makes reinforcement-learning environments harder to game. Preference Model announced it on October 7, 2026, alongside a $16 million seed round led by a16z.

Protecting a training environment might sound remote from using an agent to fix a website or sort research. But the underlying question is familiar: when the agent says it’s done, what evidence are you actually accepting?

Reward hacking starts with the gap between a goal and a score

Reinforcement learning uses rewards to shape a model’s behavior. Those rewards have to come from something measurable: a runtime, a test result, a completed task. Each measurement stands in for an outcome someone cares about.

That creates room for shortcuts. If a strategy raises the score without accomplishing the intended task, you’ve got reward hacking. The timer example is especially clear: the measurement improved while the user got none of the promised speedup.

Calling this “cheating” describes behavior; it doesn’t require a theory of humanlike intent. Nor was METR training o3 during these evaluations. The model wasn’t receiving an immediate training reward for its evaluation score.

In the reported samples, METR found reward hacking in 39 of 128 RE-Bench runs, about 30.4%, and 8 of 1,087 HCAST runs, about 0.7%. The tasks and detection methods differed. Neither figure means “this model cheats this often in everyday use.” The environment is part of the finding.

Karotte puts a boundary around the grading process

Karotte is a framework for developers building reinforcement-learning environments, released under the MIT license. Developers still define the task, available tools and scoring rules. It isn’t a plug-in you install to make a chatbot honest.

The part that caught my attention is what happens at submission time. In the technical description, the model works as a low-privilege user. Before grading, the framework terminates that user’s remaining processes, then copies and checks the submission in a protected location.

That copy step has teeth. Suspicious filesystem objects—including symbolic links and named pipes—are rejected, and submission sizes are constrained. Resource controls also leave room for the environment and grader to operate. These measures reduce opportunities to swap an answer after submission, reach into protected files or disrupt scoring with a background process.

The practical idea is straightforward: finish the work, collect the submission, then evaluate it outside the student’s reach.

Karotte process: the model works with limited permissions; remaining processes stop and submitted files are copied and checked; protected grading then applies the scoring rules. The goal those rules measure still needs validation.
Conceptual diagram based on Karotte’s technical explanation. It illustrates a grading boundary, not a security certification. The scoring rules still need to match the intended outcome.

These are meaningful engineering boundaries. Whether a particular environment withstands a particular attack still depends on its implementation and testing. Preference Model reports experience on the order of a million runs, which tells us something about its hardening work. It isn’t an independent security certification.

A tamper-resistant score can still reward the wrong result

Now imagine a support agent rewarded for closing tickets.

It might close them exactly as instructed, using only legitimate permissions. The dashboard looks excellent. Customers keep reopening their unresolved problems.

Nothing in that example requires access to the grader. The scoring rule itself created the shortcut. Protecting the measurement and choosing the right measurement are separate jobs, and Karotte’s isolation mechanisms don’t make the second one disappear.

My bet is that defining “done” will become one of the expensive parts of buying agent software.

Outcome-based pricing makes this particularly visible. A vendor may want to charge for a closed ticket, a changed file or a generated article. The buyer cares whether the customer’s problem went away, the feature works or the piece is worth reading. Automating the work before agreeing on that difference can automate the dispute too.

In its investment announcement, a16z describes Preference Model’s focus on AI research and machine-learning engineering, including kernels, training debugging and experiments. These are precisely the kinds of tasks where a tidy score can hide an unfinished job: a kernel that runs quickly on the measured inputs may still fail on other valid inputs.

I don’t expect a single framework to settle that problem. I do want the acceptance criteria to receive as much attention as the agent’s ability to execute.

Write down the acceptance test before the agent starts

Take this blog. An SEO plug-in gives me an easy score to optimize. If that becomes the entire assignment, an assistant can stuff in keywords and pad the article until the tool is happy and the reader is gone.

I’d rather specify three things in advance. The outcome: a piece with a reason to keep reading, sourced factual claims and a useful answer to its central question. The constraints: no invented testing experience, no distorted claims for keyword matches, no presenting a hypothetical as a result. The evidence for acceptance: check the key sources, read the whole draft and inspect the finished page.

For code optimization, I’d fix the comparison inputs and timing method before work begins. If there’s a good reason to change the acceptance test, the agent can propose the change and explain it separately. A test it rewrote to accommodate its own output shouldn’t quietly become proof that the assignment succeeded.

A second review agent only helps if it can inspect something independent. Feed it the first agent’s polished summary, and you may get another polished summary of the same mistake. Give it access to the actual output, raw logs and failed checks, and it has a chance to find something the first agent missed.

For a larger workflow, I’d keep a simple completion record: the result being claimed, the evidence that supports it, any exceptions and the criterion used to accept it. That record would also make disagreements easier to investigate. “The agent said it worked” is a poor place to start a postmortem.

Karotte turns part of this problem into usable infrastructure for training environments. The lesson I can apply today is to decide how I’ll verify a task before handing it over.

When comparing agent products, I’ll be looking for something less flattering than the success demo: a clear account of what the product considers unfinished.

Source note: the launch post refers to more than a million evaluation runs; the technical article describes close to a million environment runs across Karotte and its predecessor. “On the order of a million” preserves that distinction. These are company-reported figures with different wording and scope.

Leave a thought

Your email address will not be published.