A $1,200 Experiment Taught a Tiny 4B Model to Beat Postgres's Default Query Plans — With Caveats Worth Reading
Developer Rohan Bansal used reinforcement learning to train a 4-billion-parameter model to rewrite database query plans. The headline 1.81x speedup is real but comes from picking the best of up to 15 attempts on a well-known benchmark. The model's own choices averaged 1.40x. Here's what was done, what it cost, and why small specialist models are the point.

The takeaway: Independent developer Rohan Bansal has published a detailed write-up of training a small, 4-billion-parameter language model to improve the query plans chosen by Postgres, one of the most widely used databases. On a standard benchmark of join-heavy queries, the best plans his model found (picking the best of up to 15 attempts per query) ran 1.81 times faster than Postgres's own plans, cutting total run time by 44.7%. He says the whole project cost about $1,200. The result is interesting less for the number than for what it shows: a cheap, specialised model trained on a verifiable task can learn to do that one job well.
The problem: query planners guess
When you run a database query, the planner has to decide things like the order to join tables and which indexes to use. Bansal points out that choosing join order is NP-hard, so planners rely on estimates, and those estimates can be badly wrong. What makes this a good fit for AI training is that checking a plan is easy: you run it and time it. In his words, "Language models are particularly good at learning how to do tasks with easily verifiable outputs."
What he built
- The model: Empero's Qwen3.8-4B-Distill, a Qwen 3.5 4B model distilled from the 2.4-trillion-parameter Qwen 3.8 by a small German lab, small enough to run on his home rig of two RTX 3090 GPUs.
- The harness: an agent called qo-agent that can inspect tables and statistics, look at Postgres's default plan, and test up to five candidate plans before choosing one or keeping the default.
- Training in two stages: first, fine-tuning on example runs generated by a large frontier model (OpenAI's GPT-6 Astra); then reinforcement learning, where the model's plans were timed against Postgres and faster plans were rewarded.
- The data: the Cardinality Estimation Benchmark (CEB) for training, and the Join Order Benchmark (JOB) for testing. Both use the IMDb dataset. He checked that no training query had the same join structure as any test query, so the model couldn't overfit to the test set. None did, so nothing had to be removed.
The results, read carefully
On all 113 JOB queries, three rollouts each:
- The model's own final choices averaged a 1.40x geometric-mean speedup.
- Picking the best candidate across three runs, up to 15 candidate plans per query, gave 1.81x, with 68 wins and no regressions.
The 1.81x figure is the headline number and it's accurate, but it depends on that best-of-15 selection. Bansal argues this matches how someone tuning a batch of queries would actually use the model: sample several times and keep the fastest. That's reasonable, but it's worth knowing the single-run figure is lower.
Two other caveats:
- It's one database. Training and testing both used IMDb data. Bansal argues that's the point, since a company would train on its own databases, but it isn't evidence the model generalises to databases it hasn't seen.
- Timing is noisy. Much of the write-up covers how measurement noise could "fool" the reward. For one query under his initial settings, a plan identical to the default could show a phantom 14–26% speedup or slowdown about a fifth of the time. Giving Postgres more cache memory largely removed that, and it's a useful warning for anyone building similar reward signals.
What the model learned
Bansal's trace analysis shows the model mostly inspected tables, statistics or Postgres's default plan before proposing plans, used all five attempts in most searches, and leaned heavily on scan hints and join-order rewrites. On one query (job-01d) it found a 90x speedup: it spotted a slow sequential scan and forced a bitmap scan instead.
What it cost
About $800 to rent a node with two H100 GPUs for roughly 95 hours, plus about $400 in OpenAI API fees to generate the training demonstrations from Astra. Total: $1,200 in rented compute and API fees, by his account, not counting his own GPUs and electricity.
Why it matters
In our view, this is a clear example of a trend that matters more to most businesses than frontier benchmarks: a company with a narrow, checkable task and its own data can train a small open model to do that task well, cheaply. Bansal's own conclusion is that frontier models aren't going anywhere (his small model learned from one), but "small models will be increasingly used" for niche jobs "as they are faster and cheaper to train". This is a single-author experiment on one benchmark, not a product. It was published on 16 September and has not yet been independently reproduced, though he has released the code.
Ask Relay — he reads every question himself and replies personally by email.
