Overview
I trained a driving agent with PPO (Policy-Gradient Reinforcement Learning) in a Python simulation, then brought it to the web so people can race it. Each lap is qualifying rules, one lap from a standing start, and leaving the road voids the lap, for you and for the agent.
The hard part of this project wasn't training the agent. It was making sure the browser game is exactly the world the agent trained in and coding the physics of the car and the environment. If the TypeScript physics or sensors drift even slightly from Python, the agent sees inputs it has never seen and drives worse, with no error to tell you why.
Tech Stack
- Python, Gymnasium and Stable-Baselines3 (PyTorch): the simulation environment and PPO training.
- ONNX: the trained policy is exported as a single ~32 KB file with only the actor network.
- TypeScript and onnxruntime-web: the browser port of the physics and sensors, and the runtime that runs the model every few frames.
- Next.js, React and SVG: the game page, drawn as an SVG with a follow camera.
- Vitest: parity tests that replay Python's recorded laps against the TypeScript port.
Features
- Race a live agent: the blue ghost car is the model deciding in real time, 15 times a second.
- Three difficulties: earlier training checkpoints make slower drivers:
- Easy: 250k training steps, about a 25.5 s lap.
- Medium: 500k training steps, about a 15.7 s lap.
- Hard: the fully trained agent (After 2M steps), 11.06 s.
- Qualifying rules: leave the track and your lap is invalidated, the same rule the agent was trained under.
Architecture
The simulation is shared by design: one config.json holds the physics constants and one generated track file holds the walls and checkpoints. Python and TypeScript both load them, so the numbers can't drift apart.
Python sim + Gymnasium env → PPO training → ONNX export (+ parity file)
↓
Browser: TS physics + sensors → onnxruntime-web → throttle / steer
What the agent sees and does
- Observation (12 numbers): 9 distance rays fanned from -90° to +90° around the car, its speed, the angle to the next checkpoint, and whether it's off the road.
- Actions (9): every combination of throttle (brake, coast, accelerate) and steering (left, straight, right).
- Timing: physics runs at a fixed 60 Hz; the agent picks an action every 4 steps and holds it, like in training.
Reward
The agent earns reward for passing checkpoints in order and a bonus for finishing the lap, pays a small cost every step (so faster is better), and takes a large penalty for leaving the road, which also ends the episode.
How was it made?
I started by building everything in Python, with the track as my first goal. In this early version, I didn't focus on designing the track itself. Instead, I wrote a script that takes a list of points and connects them into a track with a fixed width. For the first layout, I used an AI assistant to generate the list of points, and my script turned them into the track.
Once I had the track, I moved on to the car physics. I went with a simple kinematic model and added grip, so the car turns more sharply at low speeds and less at high speeds, which adds a bit of challenge. The car's action space is just throttle, brake, and steering.
After the logic was done, I ported everything from Python to TypeScript. I added test benchmarks to check that the physics behave the same in both versions, so the agent acts the same way on the website as it did in training.
Next, I built a viewer tool so I could drive the car myself and test whether the physics felt right and were fun to drive. A config file holds the car's settings, such as top speed and turn radius.
With both of these in place, I started developing the agent.
The agent observes 12 floats: 9 rays that let it see what's ahead, plus its speed, the angle to the next checkpoint, and whether it's off the road. It then picks one of 9 actions, every combination of throttle and steering.
The agent's rewards are shown below.
| Event | Reward |
|---|---|
| Pass the next checkpoint (in order) | +1 |
| Complete a lap | +10 |
| Each agent step | -0.05 |
| Go out of bounds (voids the lap, ends the episode) | -5 |
An episode ends when the agent completes a lap, goes off track, or fails to reach a new checkpoint within 5 seconds. As of 27 September 2026, only one version of the car has been tested, achieving a lap time of 11.06 seconds after 2 million training steps.
Challenges & Solutions
Keeping two simulations identical. The export tool records one full deterministic lap per model: the car state, the observation, the network's output and the chosen action at every step. A Vitest suite replays that lap through the TypeScript physics and the ONNX model and checks every step against Python, so a physics or sensor bug fails a test instead of silently making the agent worse.
Running a model inside a game loop. Physics steps are synchronous but inference in the browser is asynchronous. The ghost car pauses its physics while it waits for a decision and then catches up, so it still gets exactly 4 physics steps per decision, and its lap time in the browser matches Python's to the hundredth.
Keeping the download small. The model is tiny, but the ONNX runtime isn't. Using the CPU-only WebAssembly build of onnxruntime-web halved the runtime download, and all three difficulty models together add under 100 KB.
Future Plans
Track Creator: Let users create their own tracks.
Easier website integration: Currently, too many imports are needed to replicate the game's behavior on the website.
Other models and hyperparameters: So far, only one model has been trained. We can experiment with other hyperparameters to see how the agent improves.
More complex car system (maybe): This would add more depth to the game, but for now, keeping it simple is better.