FernFern← Home

Whitepaper · Updated September 2026

High-fidelity environments for evaluating and training robot policies

Fern is building scalable evaluation and reinforcement-learning environments for robotics. We build each environment from real robot data, so any policy can be benchmarked and improved without ever running on physical hardware.

The problem with evaluating robot policies

General robot policies are improving fast — but the way we measure them hasn’t. The dominant evaluation methodology is still “put it on a real robot, run a few rollouts, eyeball the results.” That has several problems.

  • It doesn’t scale. Each evaluation costs operator time, hardware wear, and resets between trials — meaningfully testing a single checkpoint takes hours, not seconds.
  • There’s no apples-to-apples comparison. Two checkpoints are never run under identical conditions, so a difference in success rate could just as easily come from drift in the setup as from the policy itself.
  • It isn’t reproducible. Lighting, object placements, and contact dynamics drift between sessions, so results don’t transfer cleanly across runs or between research groups.
  • RL is impossible on hardware. Reinforcement learning only produces a useful signal when you can run thousands of rollouts in parallel at scale — and there is simply no way to do that on physical robots, where every rollout ties up a real machine in real time.

The same problems were solved for game-playing agents and language models with simulators and benchmark suites. Robotics doesn’t have either yet.

Why not build your own simulator?

The obvious answer is a simulator. In practice, running one in-house is expensive:

  • Closing the sim-to-real gap takes days of tuning. Geometry, cameras, contact parameters, and controller behavior all have to be matched to the real robot before a policy behaves the same in sim as it does in production.
  • Maintaining a sim is a job in itself. Every new object, fixture, or cell layout means more tuning that pulls engineers away from the policy.
  • Scaling out is its own infrastructure project. One sim on a workstation is easy. Thousands running in parallel means clusters, scheduling, rendering, and logging, and most teams never get past a handful of rollouts at a time.

Fern takes all of that off your plate:

  • Massive parallel scale. We run thousands of sims in parallel, so an eval sweep that would tie up your robots for weeks finishes in hours, and RL gets the volume of rollouts it actually needs. Scale is what makes evaluation and RL effective.
  • Zero real-to-sim gap. We guarantee the sim matches your production environment exactly, so your robot policies run in it just as they do on your robots.
  • No infrastructure to run. We build, host, and maintain the sims and the compute behind them. Your team submits policies and gets results back.

The gap closes in the other direction too: policies trained in our sims have transferred to our clients’ real robots with very little sim-to-real gap.

We know what that takes because we’ve worked on the policy side. We’re PhD researchers and AI engineers who have trained robot policies for our clients, so we know what a sim has to get right for a policy to behave the same in sim and on real hardware.

Let Fern handle the sim

Building and maintaining high-fidelity sims is a full-time job, and it isn’t the job your team was hired to do. Fern builds the sim for your robots and production environment, keeps it matched as your setup changes, and runs thousands of copies in parallel for your evals and RL training.

Your robotics researchers and engineers get that time back to focus on what actually moves the needle: improving the policy.

Want to stop maintaining sims and start shipping better policies? Reach out at founders@fern.bot or book a call.