Docs
SnakeBench measures how well a language model plans a long sequence of actions with no feedback. The model sees a whole Snake game up front, the board, the walls, its body and the full ordered list of food that will ever spawn, and must answer with one movement string. The simulator then plays that string exactly as written.
It tests long-horizon planning, spatial reasoning, state tracking and optimization at once.
Benchmark flow
- Configure. Choose grid size, obstacle count, food count and seed. A live preview shows the exact board.
- Prompt. Copy the generated prompt into any LLM, then paste its answer back.
- Simulate. Watch the plan play out at 1×, 2× or 5×, or skip straight to the end.
- Score. See food collected, survival, efficiency and how it stacks up against the reference solver.
Board and coordinates
- Cells are
(x,y).(0,0)is the top-left.xgrows right,ygrows down. - The snake starts with length 3 at the centre column, head at
(⌊n/2⌋, ⌊n/2⌋), body trailing downward, facing up. - Obstacles never sit on the snake or the two cells straight ahead of it, and never cut off part of the board.
- Obstacles are capped at 20% of cells.
Command syntax
A response is a sequence of tokens:
N, a positive integer: move forwardNcells.L/R: turn 90° left or right, relative to the current facing. Turning does not move the snake.
Example: 6L1R2L1L6 means forward 6, turn left, forward 1, turn right, forward 2,
and so on. Whitespace and backticks are stripped and letters are case-insensitive. Anything
else, or a 0, makes the response invalid.
Rules and edge cases
- One food at a time. Only the current food is on the board. The next appears the moment it is eaten.
- Growth. Eating grows the snake by 1: the tail stays put on that move.
- Tail chasing. The tail vacates its cell during a move, so moving into the tail's cell is safe, except on a move that eats.
- Food placement. Food never spawns on an obstacle or inside a pocket the board could only reach through a single cell. It can spawn under the snake's body, and is only eaten when the head enters that cell. Consecutive foods never share a cell.
- Death. Leaving the grid, entering an obstacle, or entering your own body ends the game immediately. The fatal move does not count.
- Running out. The game also ends when commands are exhausted, all food is eaten, or the move limit (
food × grid × 4) is reached. - Malformed responses score 0 food and an F. Nothing is guessed or repaired.
- Trailing turns after the last move are allowed and do nothing.
Scoring
- Food collected is the primary score.
- Moves survived counts every successful move.
- Efficiency is the sum of Manhattan distances between consecutive food pickups divided by the moves actually used to reach the last one. 100% means no detours at all, which walls and the body rarely allow.
- vs reference compares food collected to a built-in solver: BFS to each food, only taking paths that keep its tail reachable, otherwise following its tail until a safe path opens. It is a solid baseline, not an optimal player.
| Grade | Meaning |
|---|---|
| S | Ate every food |
| A | ≥ 90% of the reference solver's food |
| B | ≥ 70% |
| C | ≥ 50% |
| D | ≥ 25% |
| F | Below 25%, or an invalid response |
Difficulty tiers
Tiers are presets. Report results per tier, averaged across several seeds.
| Tier | Grid | Obstacles | Food |
|---|---|---|---|
| easy | 18×18 | 36 | 35 |
| medium | 24×24 | 90 | 60 |
| hard | 28×28 | 140 | 80 |
| brutal | 32×32 | 204 | 100 |
Determinism
Boards and food come from a seeded mulberry32 generator, so the same grid size,
obstacle count, food count and seed always build the same game in every browser. The
simulator has no randomness. To compare models fairly, give each one the same settings and
seed, in a fresh conversation, with the prompt unchanged.