Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback
Singla, J., Pareek, P., Jawanpuria, P., & Singla, P.
Submitted to ICLR 2027, arXiv:2609.37356
Summary: Benchmarking a mixed-integer linear programming (MILP) solver needs feasible, hard instances, but most generators do not learn hardness; they inherit it from seed instances or search for it per family and size. Here a frozen SCIP solver scores each instance a language model writes by branch-and-bound nodes and the relaxation gap left after root cuts, so weaknesses the solver repairs earn little, while a validity gate rejects hardness bought with unbounded continuous variables, aggregated big-M links or extreme coefficient ranges. Trained by reinforcement learning without seed instances, OptiScribe-12B raises median nodes over base Gemma-4-12B 1.7–5.0× on capacitated facility location (CFL) and 1.9–4.5× on max-cut, and CFL feasibility by 9–19 points; the gains hold beyond training sizes and under held-out solvers, HiGHS and Gurobi. Multiple knapsack stays easy for every model.