The article walks readers through creating a “decision model” that forces a language model to choose only among a predefined set of answers, such as A through E. It begins by explaining that standard models generate text token by token, but for structured output one can constrain the vocabulary so the model can emit only the tokens representing the allowed options. By masking all other tokens in the model’s output layer, the highest‑probability remaining token becomes the prediction.

A concrete example uses the Qwen/Qwen3‑1.7B model. The provided Python script loads the tokenizer and model, builds a prompt that lists a question and each option, and then extracts the logits for the final token position. After applying a softmax only to the logits corresponding to the option tokens, the script prints the predicted letter and the probability for each choice. When run on a simple question—“What color is the sky?”—with options Red, Blue, Green, Purple, and I don’t know, the model returns B (Blue) with a probability of 0.9988, while all other options receive near‑zero scores.
To evaluate broader performance, the author tested the script on a random holdout from the CommonsenseQA dataset. Before any fine‑tuning, the model achieved an accuracy of 725 out of 1221 examples (≈59.4 %) and a macro‑averaged F1 of 0.584. After a quick fine‑tune on the same dataset, accuracy rose to 762/1221 (≈62.4 %) and macro‑averaged F1 to 0.623.
The piece then highlights a calibration problem: when the model is very confident (probability >0.9), it is correct only about 70 % of the time, indicating overconfidence. A reliability diagram shows that predictions in the 0.9‑1.0 confidence bin are far less accurate than the confidence suggests.
To improve calibration, the author applied temperature scaling. By searching for a temperature value that better aligns confidence with observed accuracy, they found a factor of approximately 3.797. After re‑scaling the logits with this temperature, the confidence bins more closely match accuracy—for instance, the 0.9‑1.0 bin now shows ~93 % confidence with ~95 % accuracy.
The post concludes by linking to a GitHub repository that contains scripts for building a dataset, evaluating, fine‑tuning, and calibrating a decision model, encouraging readers to adapt the approach to larger models or other tasks.
