Choose the action, then the network
The useful split is how fine the action is. A high-level agent picks a final placement: column, rotation, and optionally a hold. A legal-move generator lists those landings, usually a few dozen, and the network only has to rank them. A low-level agent presses left, right, rotate, and drop each frame. That is closer to the keyboard and much harder to learn.
State and reward
The board can be the raw grid, or a short feature vector: column heights, holes, bumpiness, wells, the current piece, and the next piece in the queue. A sparse reward is lines cleared, with a small cost per piece so survival is not free. Shaped penalties for holes and height help early training and can also teach the agent to play for the penalty instead of the clear.
Variant 1 — afterstate value from features
Enumerate each legal landing, score the board that landing would leave, and take the best score. A small multilayer perceptron, on the order of 32 to 128 units and two or three layers, predicts that value. Training can start by copying a hand-written heuristic, then fine-tune with temporal-difference updates. This is fast on a CPU and only as good as the features.
The published game is this shape. Ten features describe the placement. A network with hidden layers of 16 and 12 writes ten weights for those features. Finished games stay in this browser, up to the last 1000, and training uses that history. Nothing is uploaded.
Variant 2 — actor-critic on the grid
A convolutional network reads the board as channels and outputs a policy over placement bins plus a value. Invalid landings are masked. PPO or A2C with an entropy bonus can learn T-spin and cavity patterns that a fixed feature list misses. It costs more per move than the feature model, and the action index has to stay aligned with the move generator.
Variant 3 — policy, value, and search
An AlphaZero-style net proposes a policy and a value. Monte Carlo tree search uses that value at the leaves and returns a stronger policy than the network alone. Self-play stores the improved policy and the game outcome, and the net is trained to match both. This plans further ahead and is the slowest of the five: a move can take tens of milliseconds once the simulation count rises.
Variant 4 — a recurrent agent and the queue
A recurrent layer keeps context across pieces, so the agent can time a hold or a setup that needs the next piece in the queue. The cost is credit assignment: a good clear may be the result of several earlier decisions, and the training sequences have to be packed with that in mind.
Variant 5 — evolve the weights
Treat the weights of a linear heuristic, or of a one- or two-layer network on the same features as variant 1, as a genome. CMA-ES or another evolutionary strategy samples weight vectors, plays them, and keeps the ones that clear more lines. There is no backpropagation. Evaluation is expensive because each candidate has to play.
Which one to build
- Use variant 1 when the agent has to move inside a page.
- Use variant 2 when the pattern lives in the grid and features are not enough.
- Use variant 3 when a slower move is acceptable and lookahead matters.
- Use variant 4 when the next-piece queue is part of the plan.
- Use variant 5 when you want to search weights without a gradient.
A practical curriculum starts with slow gravity and short games, adds combos and faster drops, then measures lines per game, holes, bumpiness, and milliseconds per move.
Source
The game and this design note are in the Tetris repository on GitHub, published under the Unlicense. It is part of the open-source tools Polymech publishes while building manufacturing software.