Small-Model Linguistic Reasoning under Compute Constraints
Conference on Computational Linguistics and Speech Processing (ROCLING 2026) 2026
Linguistics Olympiad problems require inferring linguistic rules from a few examples. We study which training and inference methods help a 4B open-weights model solve such problems under the compute constraints of the IOL-AI 2026 challenge, which allows one T4 GPU and 30 minutes. We generated solutions to 175 problems with a frontier LLM; former national-team contestants reviewed 150 and flagged flawed reasoning far more often than wrong answers, revealing a gap between answer correctness and acceptable linguistic reasoning. Trained on full-length solutions, the model could not finish its answers within the generation limit, while a model trained on shortened solutions scored similarly to full-length reasoning at a fraction of the length. We also tested multilingual and in-domain training data, voting and three variants of it, and grammar retrieval; none improved consistently across evaluations. Our 4B system placed fourth of 48 teams, behind three teams that all used a 14B model.
