Reinforcement-learning bots on a Roblox game server
October 1, 2026
Hyperdrive, our car-soccer game on Roblox, runs reinforcement-learning bots on its game server. They are Necto and Nexto, two open-source Rocket League bots trained with self-play, ported to Luau. They play against real players at nine skill tiers, and when a player leaves a match, a bot matched to that player's rating takes the empty seat.
A simulator that matches the training data
Policies trained in one physics engine usually fall apart in another. We avoided most of that by building Hyperdrive's car and ball simulation on the conventions of RocketSim, an open-source reimplementation of Rocket League physics: 120 physics ticks per second, z-up, and an internal unit of one fiftieth of a Rocket League unit. The server runs two physics ticks per 60 Hz frame.
Because the units line up, the observation builders read the simulation state directly, multiply by 50, and normalize exactly as the training code did. There is no coordinate remapping layer to get wrong. Nexto's attention layers do not care about the order of entities, so the server can list cars in whatever order its slots happen to be in.
The port
Roblox has no file storage for model weights, so each model's float32 weights are base64 text split across ModuleScripts of about 300 KB each. At startup the server joins the chunks and decodes them into one buffer.
| Model | Architecture | Weights |
|---|---|---|
| Nexto | attention, 2 blocks, width 128, 4 heads, 90 actions | 444,752 floats, 1.78 MB |
| Necto | attention, 1 block, 5 action heads | 173,964 floats, 0.70 MB |
Each model ships with a golden input and output exported from PyTorch, and a self-test that runs the Luau forward pass against it. Nexto matches to a maximum difference of 2.1e-6 and picks the same action. Necto matches to 2.9e-6.
The first Nexto port took 6.4 ms per forward pass. A native-compiled kernel that unrolls four output rows per loop brought that to 1.27 ms. Necto takes 0.85 ms. Each bot decides every 8 physics ticks (15 times per second) and holds its action in between, and the server staggers decisions by slot so a full 3v3 lobby never runs six forward passes in the same frame.
A skill ladder from one network
Both networks output a distribution over actions. At full strength the bot takes the most likely action. For weaker tiers it samples from the logits with a temperature set by a parameter β, using a scale of ln((1 + β) / (1 − β)) / ln 3. Lower β flattens the distribution and the bot makes more mistakes. Nothing else changes between tiers: no added noise, no slower reactions.
| Tier | Model | β | Rating seed |
|---|---|---|---|
| God | Nexto | 0.85 | 2450 |
| SSS | Nexto | 0.70 | 2200 |
| SS | Nexto | 0.55 | 1900 |
| S | Nexto | 0.40 | 1550 |
| A | Necto | 0.85 | 1100 |
| B | Necto | 0.70 | 800 |
| C | Necto | 0.55 | 500 |
| D | Necto | 0.35 | 250 |
| E | scripted | 0.30 | 0 |
Each tier carries a rating on the same scale as players. A new player at the default rating of 1400 is expected to beat tier E 96.5% of the time and tier God 0.8% of the time. Matchmaking uses those numbers to pick the bot for a backfill seat or a bot-filled casual match, aiming for a target win rate: 70% for new players, 50% normally, 47.5% for a player on a win streak.
What broke
Abilities
Hyperdrive has abilities that Rocket League does not, such as a grappling hook and a ball freeze, and some of them push cars well past normal speeds. The bots never saw any of them in training. The observation code normalizes values by Rocket League's limits but, like the upstream code, never clamps them, so a car moving faster than 2,300 units per second produced inputs far outside anything the network had seen, and the attention layers returned nonsense. We now clamp every observation to the training limits before normalizing: car speed 2,300, ball speed 6,000, angular speed 5.5 for cars and 6 for the ball, and positions to the field bounds. Abilities also had to write velocity into the simulation state the bots read. A recovery controller turns bots upright after an ability flings them.
Team play
We also ported three stronger community bots: Element, Immortal, and Karma, all plain multilayer perceptrons (Karma is 4.5 million parameters, 18 MB of weights across 80 chunks). All three were trained for 1v1 only, and in team modes the copies tended to chase the ball together. We built a role layer that let one bot per team run its network while teammates held scripted support positions. It passed our harness, but it replaced much of the networks' play with scripted positioning, so it is off by default. Those three models are switched off in live play, and the ladder runs on Necto and Nexto.
Kickoffs
In early team matches, bots sometimes took more than 30 seconds to touch the ball after a kickoff. After fixes to kickoff detection and staging, first touch lands in 1.7 to 2.7 seconds. One fix was for the ball freeze ability: a frozen ball near midfield looked like a kickoff to the old detector, so the server now arms kickoff detection only during the actual kickoff countdown.
Motion
A bot holds each action for 8 ticks and then switches, which looks twitchy next to a human on an analog stick. An optional smoothing step blends the steering and throttle toward the new action over a few frames. Buttons such as jump still switch instantly.
Testing bots without watching them
Most of these problems were reported as “the bots feel broken,” which is hard to fix. We wrote a harness that runs staged matches on the server and measures: time to first touch after each kickoff, how often each team's lead attacker changes, time spent with no one playing the ball, circling, cars stuck in place, invalid physics states, and the possession split. A configuration passes when the median of three runs stays inside the limits, including first touch within 2.8 seconds on every kickoff and possession between 35% and 65%. The last certified run had a worst kickoff of 1.77 seconds.
Credits
Necto and Nexto are by Rolv-Arild and contributors to the RLGym project, and we use them with the author's permission. Element is by Rangler, Immortal by CosmicVivacity, and Karma by Kuenec. The bots were trained with RLGym, and Hyperdrive's physics follows RocketSim.