A language model that runs entirely inside Roblox
October 1, 2026
We run Qwen3-0.6B, an open language model with 0.6 billion parameters, entirely inside Roblox. The forward pass is Luau. The weights live on the platform as 144 image assets. Nothing calls out to an external server at runtime, and on our test prompts the model produces the same tokens as the PyTorch reference.
Why put the model in the engine
Our goal is agents that build, test, and play Roblox games. Most of that work happens inside the engine: reading the scene graph, editing scripts, calling engine APIs, talking to players in a live server. A model hosted elsewhere reaches all of that through HTTP, and every live game that uses it depends on someone else's GPU staying up.
A model that runs in Luau has none of those dependencies, and building one tells us how much inference the Roblox runtime can do.
The port
The weights are quantized to int8 with one scale per row. Activations stay in 32-bit floats. The Luau side has a small tensor module (int8 matrix-vector product, RMSNorm, rotary embeddings, SiLU, attention), a KV cache, a byte-level BPE tokenizer with the full 151k vocabulary, and the Qwen3 chat template. The template renders byte for byte the same as the Hugging Face implementation, which matters because one wrong whitespace token is enough to derail tool calling.
The matrix kernel is register-blocked: each inner loop reads one int8 row and keeps several partial sums in locals. We use the same kernel style for the neural bots in our car-soccer game.
Against the PyTorch int8 reference:
- maximum hidden-state difference: 4.9e-4
- top token agrees at 34 of 34 positions
- greedy generation is identical
Cost per token
| Measurement | Value |
|---|---|
| int8 kernel throughput, interpreted Luau, one core | ~1.1 GMAC/s |
| one token, interpreted (Lune) | 0.55 s |
| prefill, Roblox Studio with native code generation | 0.128 s/token |
| decode, Roblox Studio with native code generation | 0.174 s/token |
| weights, int8 | 600 MB |
| per decoder layer / tied embedding | 15.7 MB / 155 MB |
Batching prompt tokens does not raise throughput. The kernel is bound by scalar arithmetic, so prefill costs about the same per token as decode. The tied output projection (151,936 by 1,024) accounts for roughly a quarter of the time per token.
Storing weights as images
Roblox has no file storage for a 600 MB blob. A script tops out at 200,000 characters, a single buffer at 1 GiB, a DataStore key at 4 MB. HttpService can stream the bytes, but that brings the external server back.
We store them as images. We pack the weight bytes into 144 RGBA images of 1024 by 1024 pixels, 4 MB each, and upload them through Open Cloud. At load time the game opens each one with EditableImage and reads it back with ReadPixelsBuffer. The read is byte-exact, alpha channel included, and every chunk carries a 32-bit checksum that the loader verifies before the bytes go into the weight buffers.
A cold load of all 144 images takes about 230 seconds in Studio. On a game server with the images cached it takes about 17. EditableImage only opens images owned by the experience's creator, so the uploads and the game must belong to the same account or group.
Running it on a game server
A Roblox server that blocks for a fifth of a second per token stops simulating the game. The model yields instead: inference runs in 12 ms slices per Heartbeat, and the server holds 60 frames per second while it generates.
The system prompt and tool schemas are the same for every request, so the server encodes them once at startup and keeps their KV cache. To speed that warm-up we split the model across 7 Actors with 4 layers each and stream 16-token batches through the stages. A 664-token prefix warms in 43 to 50 seconds this way, against 132 seconds on one thread. Studio gives a script 8 parallel workers; live servers give 2 or 3, so the speedup there is smaller.
Players type a request into a chat panel. The agent answers with tool calls against a small, bounded set: spawn a part, list or remove your builds, set the time of day, teleport. Each player has their own build folder, radius, and part count limit. Replies stream back token by token. With the prefix warm, a request places its part after about 25 seconds and finishes its reply in 43 to 48.

The same model also runs as a Studio plugin. There it gets tools that mirror Roblox's Studio MCP server: search the game tree, inspect an instance, read and write scripts, run Luau, read the output log.
Where 0.6B falls short
- Its knowledge of the Roblox API is thin. Asked to mark a position, it called an
add_marker()function that does not exist. - It reliably makes one tool call per request. When a request needs two steps, it often claims to have done the second without calling anything. The server now splits requests on “then” and runs the parts in sequence.
- Forty seconds per request is fine for a building toy and far too slow for anything that reacts to gameplay.
Qwen3-1.7B would know more and cost about three times as much per token. That trade does not fix the reaction-time problem, so we are splitting the job instead.
Next
The language model stays as the planner. Under it goes a small decision network distilled from a larger teacher model that labels game states. A network that size runs in about a millisecond in Luau, which is fast enough to act every frame. The planner decides what to build or where to go; the small network handles moment-to-moment control.