Article · September 2026
World models vs. physics simulation
There are two broad ways to build a simulation environment for a robot: learn one from data, or reconstruct one with a physics engine. We’ve built both. Here’s how they compare, and how we choose between them.
Learned world models
A neural network trained on the robot’s camera footage and action logs, which predicts the next camera frames given the current frames and an action.
- Strength: no tuning. The model is trained directly on your robot’s recordings. There are no assets to build and no physical parameters to fit.
- Strength: matches the data exactly. The output looks like your cameras and behaves like your recordings, including materials like deformables and granular media that are hard to model by hand.
- Limitation: limited to shorter horizons. World models are autoregressive: each frame is predicted from the model’s own previous frames, so small errors compound over a rollout. That limits them to shorter-horizon tasks.
- Limitation: can’t generalize or swap things out. The model only knows what’s in its training data. Changing an object, the lighting, or the layout means collecting new data and retraining.
Physics simulation
An explicit model of the robot, scene, and objects: geometry, cameras, and physical parameters, stepped forward by a physics engine.
- Strength: generalizes. Assets, lighting, dimensions, and layouts are explicit, so you can swap them out, randomize them, or test scenes your fleet has never seen, all without retraining anything.
- Strength: any horizon, no degradation. Physics isn’t predicted from previous frames, so rollouts stay stable for as long as the task needs.
- Strength: exact state. Every object pose and contact is known, so rewards and success checks are precise and cheap to compute.
- Limitation: needs tuning. Geometry, cameras, contact properties, and controller behavior all have to be tuned against the real robot before the sim matches reality. That’s the work Fern does for you.
How we choose
Fern’s platform is implementation-agnostic. We don’t sell a world model or a physics engine. We sell the environment: the highest-fidelity environment we can build for your task and your needs. Long-horizon tasks, and anything where you want to vary the scene, point to physics simulation. Short-horizon tasks on a fixed setup can favor a learned world model. Whichever we use, what you get stays the same:
- The same interface. Your policy runs closed-loop, unchanged, against a Gym-style environment for both evaluation and RL.
- Your production environment. Your embodiment, your scenes, your cameras, not a generic stand-in.
- Managed at scale. We run rollouts massively in parallel and handle all of the infrastructure.
Example: a learned world model
Our first environments were learned world models, and they back the results in our customer case studies. Below, the left half is a real robot recording and the right half is the model’s rollout from the same starting frame and the same actions, on an episode it never saw in training.
rice scooping · production bimanual setup
Case studyClosing the loop on a customer’s production policy. We ran their rice-scooping policy fully closed-loop in our sim and the predicted joint actions correlated exactly with the policy running on the real robot. The generated frames were also highly correlated with the real frames the robot saw in the real world. Read the case study →
Example: a physics simulation
This pumpkin-seed scooping environment was reconstructed from a customer’s own teleop recordings: the robot, cameras, pan, scoop, and all 12,900 seeds. Below, the left half is the real recording of the episode the scene was built from. The right half replays the scooping arm’s recorded actions from that episode through the latest version of the environment on our platform, with the other arm holding the pan as it does in the recording. Every seed, contact, and camera frame is simulated.
pumpkin seed scooping · recorded actions replayed in physics
Because the scene is explicit, changing it is a setting, not a new model. The customer can dim or brighten the lighting, fill the pan anywhere from empty to all 12,900 seeds or clear out specific areas of it, and swap the scoop for a different tool, like a measuring cup, then evaluate the same policy against each variation on the next reset. A world model can’t do this: it only knows the scenes in its training data, so each of those changes would mean collecting new recordings and retraining.