We needed to map a free-form user question to one of about 1,000 database table names, fast and on our own hardware. We tried Laya, the open-source “System-1 decision model,” end to end: zero-shot, with hybrid retrieval, and finally fine-tuned on synthetic data. Here is what worked, what quietly failed, and what we would do differently.
The problem
Our data platform exposes roughly 1,000 logical tables. Users ask questions like “what were the total daily sales per store in January, day by day?” and somewhere upstream of any SQL generation, something has to decide that this question is about XQR7_DTV_AGG and not about the forty other tables that also mention stores, days, or totals. A large LLM does this well, but we wanted the lookup to be a cheap, local primitive: called on every user turn, adding tens of milliseconds, not seconds, and never shipping schema metadata to an external API.
That framing — classification over a large, closed label set — is exactly what Laya advertises. Laya is a 421M-parameter, non-autoregressive decision model built on a ModernBERT-style encoder, released under Apache 2.0. You give it a state (the user’s text) and a typed question with candidate answers; it returns a choice with calibrated probabilities in a single forward pass, around 33 ms on modest GPUs. The label set is defined at request time, so no retraining is needed when tables are added. On paper, a perfect fit.
The first wall: it is not a long-context model
The naive plan — “feed the model everything about all tables as context” — dies immediately. Our per-table documentation (name, human label, description, column list, typical usage phrases, disambiguation notes) compacts to about 8.5 MB of text, close to two million tokens. Laya’s input budget is 512 tokens (1,024 for some checkpoints), of which roughly 200 are reserved for rendering the answer options, and the documentation itself notes quality drops past about 20 options per question.
So the architecture has to be two-stage:
- Retrieval narrows about 1,000 tables to the top 15 candidates for the query.
- Decision: Laya picks one candidate, given a ~140-character description of each.
This is a standard retrieve-then-rerank pattern; the interesting part is what each stage turned out to contribute.
Measuring anything at all
We built the evaluation set from an existing library of scripted support conversations: 115 turns where the expected answer text names the table(s) actually used. Two details mattered more than any model choice:
- Gold labels are noisy. The “expected” table is the one a reference agent happened to use; a miss is sometimes a defensible alternative. Absolute numbers below are pessimistic.
- Follow-up turns are unanswerable without history. “Show me the volumes for that month” retrieves nothing by itself. Concatenating the two previous user turns onto the query (
q_ctx) roughly doubled every metric in the study. If your users ask follow-ups, conversation context is not an optimization — it is table stakes.
Stage one: retrieval is the ceiling
Whatever the decision stage does, it cannot pick a table that retrieval failed to surface, so hit@15 is a hard ceiling on end-to-end accuracy. We compared three retrievers, all over the same compacted TSV (one row per table, weighted fields):
| Retriever | hit@15, bare query | hit@15, with context |
|---|---|---|
| BM25, field-weighted | 0.24 | 0.45 |
| BM25 + hand-written abbreviation aliases | 0.24 | 0.46 |
| Hybrid: BM25 + small embedder, reciprocal-rank fusion | 0.39 | 0.65 |
Share of eval turns where a correct table appears in the top-15 candidates. Embedder: a 33M-parameter sentence encoder, one vector per table.
Two findings here:
- The vocabulary gap is the dominant failure. Users say “daily sales per store”; the table is named
XQR7_DTV_AGGand its description says “aggregated transaction volumes.” Lexical search cannot bridge naming conventions built from abbreviations. A small dense embedder closed most of that gap for the cost of a 1,000×384 matrix. - Hand-curated alias lists are a trap. We wrote ~30 expansions (the abbreviation for “store” → store, and so on) and gained one point. The long tail of naming conventions is longer than anyone’s patience; embeddings learn it for free.
Stage two: what the decision model actually adds
With hybrid retrieval and context, the gold table is in the candidate list 65% of the time. Zero-shot Laya then picks it correctly in 20% of all turns:
| Configuration (with context) | Top-1 accuracy |
|---|---|
| Hybrid retrieval alone, take rank 1 | 0.17 |
| Dense embedder alone, take rank 1 | 0.22 |
| Hybrid top-15 → Laya choice | 0.20 |
| Hybrid top-50 → Laya tournament (batches of 13, winners advance) | 0.19 |
End-to-end accuracy on 115 real conversation turns; gold table counted correct if the prediction matches any table named in the reference answer.
Read charitably: within its candidate list, Laya converts about 31% — more than four times the 6.7% a random pick over 15 options would give, from 140-character descriptions it has never seen before, in one forward pass. The mechanism works.
Read honestly: the whole pipeline loses to the plain dense retriever’s first result. And the tournament scheme — feeding it 50 candidates in batches, advancing winners — made things worse, because a batch winner is only locally best and good candidates get eliminated by strong distractors early.
One genuinely useful property survived: confidence separates. Median reported confidence was 0.88–0.90 on correct picks versus 0.68–0.73 on wrong ones (with the caveat that the shipped calibration temperatures do not cover the 11+-option bucket, so treat the values as relative). That makes Laya viable as a cheap escalation gate: answer locally when confident, hand off to a large model otherwise.
The fine-tuning attempt, or: how to win the wrong benchmark
Laya’s weights are open, and the authors publish a complete fine-tuning recipe (policy-gradient training against proper-scoring-rule rewards, with a calibration pass at the end). Their own specialized checkpoint jumps from near-chance to 0.77 on their benchmark. The pitch is precisely that you can specialize it to your decision. So we did.
Training data, for free: our per-table docs include “typical usage phrases” — short strings like “favourite delivery routes, saved route combinations, preferred carrier paths.” Split on commas, that yields ~25,000 (query, table) pairs across all 1,000 tables. For each pair we built a 15-way choice question whose distractors were the retriever’s actual top candidates for that query — hard negatives matching the deployment distribution. We held out 5% of tables entirely as validation, kept the 115 real conversation turns as a test set that training never touches, and removed the usage phrases from the option descriptions so the model could not learn string matching.
The run itself: for hardware reasons the training ran on CPU — a 56-thread Xeon with the trainer patched from NCCL/CUDA to Gloo, bf16 autocast, minimum process priority. Throughput was about 20 seconds per micro-batch, roughly 17 hours per epoch, just under three days for four epochs. Not recommended, but it proves the recipe needs nothing exotic.
The result is the most instructive table in this write-up:
| Checkpoint | Synthetic val (usage-phrase queries) | Real conversations (115 turns) |
|---|---|---|
| Base Laya, zero-shot | 0.244 | 0.200 |
| After 1 epoch | 0.340 | 0.165 |
| After 4 epochs | 0.380 | 0.139 |
Fine-tuning on synthetic queries: validation accuracy up 56%, real-world accuracy down 30%. Monotonically, in both directions.
The model learned exactly what we taught it: to classify terse, keyword-dense noun phrases lifted from documentation. Real users ask long, contextual, conversational questions. Those are different input distributions, and every epoch of training pulled the model further toward the synthetic one. The held-out-tables validation set — which we built specifically to catch overfitting — caught nothing, because it was drawn from the same synthetic distribution as the training set. It validated the wrong thing, confidently.
The lesson we would put on a poster: a validation set proves generalization only across the axis on which it differs from training. Ours differed by table and matched by query style, so it could only ever measure the former. If your real traffic differs from your training data in style, your validation set must too — even if that means it is fifty hand-written examples instead of a thousand generated ones.
Caveats, collected
- Laya is not a context model. 512–1,024 input tokens, ~15–20 usable options. Any large-label-set task becomes retrieve-then-decide, and retrieval quality, not the decision model, sets your ceiling.
- Lexical retrieval fails on schema vocabulary. Abbreviated table names versus natural phrasing is the single biggest source of misses. A small dense embedder is the cheapest large win available.
- Conversation context doubles accuracy. Resolve follow-ups by concatenating recent user turns before retrieval.
- Confidence is calibrated only for small option counts. The shipped temperature table stops at 11+ options; above that, use confidence as a relative signal for escalation, not as a probability.
- Tournaments over many candidates degrade accuracy. Batch winners are locally best; widen retrieval instead of cascading decisions.
- Fine-tuning on synthetic metadata-derived queries is actively harmful if real queries look different — and a synthetic validation set will not warn you.
- Gold labels from reference answers are noisy; treat absolute accuracy as a lower bound and watch deltas between configurations instead.
Is it still worth improving?
We think yes, along three lines, in order of expected return:
1. Fix the training distribution, not the model. The fine-tuning machinery demonstrably works — it moved validation accuracy by 14 points in three days on a CPU. It optimized the wrong target because we fed it the wrong queries. The fix is to generate training questions that look like real ones: have a capable LLM rewrite each usage phrase into full conversational questions, prepend plausible dialogue context to a share of them, and mix in whatever real logged queries exist. Same recipe, same hard negatives, same holdout discipline — plus a validation slice drawn from real traffic.
2. Spend on retrieval before the decision stage. A larger embedder, embedding the query rewritten by the conversation (rather than raw concatenation), and indexing multiple views per table (name expansion, description, column list separately) all attack the 0.65 ceiling directly. Every point of hit@15 is worth more than a point of decision-stage conversion.
3. Use the decision model where it is strong: as a gate. Even zero-shot, Laya’s confidence separates right from wrong answers. A pragmatic production shape is: dense retrieval proposes, Laya adjudicates the top candidates, high confidence answers locally in ~50 ms, low confidence escalates to a large model. The small model then does not need to beat the LLM — it needs to absorb the easy majority of traffic, which is a much lower bar.
Setup for reproducibility: about 1,000 tables; per-table docs compacted to one TSV row each; BM25 with per-field weights; 33M-parameter sentence embedder with reciprocal-rank fusion (k=60); Laya 421M English checkpoint via its Python SDK; fine-tuning with the published policy-gradient recipe, 24.6k choice items, 4 epochs, effective batch 32, LR 2.5e-5 cosine; evaluation on 115 conversation turns with any-gold-match scoring. All numbers are from single runs; small deltas should be read accordingly.