Measured performance
Benchmarks
Latest full benchmark results from solverforge-bench, grouped by problem and labeled with the versions actually tested.
Latest completed full canonical nightly runs, with the versions actually tested. Quick and partial runs are excluded; these are not measurements of newer releases.
Read the gates from left to right: 1s → 10s → 60s. First, compare how many instances reach a valid solution at the earliest tested gate. Then compare quality at that same gate; later gates show improvement. A short process runtime is not a win.
Each gate ranks feasibility → earlier-gate feasibility → solution quality, never process duration. Among solvers with equal current coverage, more valid solutions at 1s wins first, then at 10s where applicable; quality breaks the remaining tie. 0% gap matches the reference. Without references, lower validated mean cost is shown instead. Quality averages cover feasible results only; different solved subsets are not a like-for-like quality comparison. Problems are compared separately.
Capacitated vehicle routing
Assign customers to vehicle routes while respecting capacity. Lower route cost is better.
CVRPLIB-X · 100 instances · 2400 recorded results · Completed
Reference values for all 100 instances, from CVRPLIB (Uchoa et al. 2017 Set X): proven optima where the source closed the instance, best known bounds otherwise.
Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.
Feasible results
Instances solved to a hard-feasible solution, out of 100.
1-second budget
- pyvrp 1s 100/100
- vroom 1s 100/100
- ortools 1s 100/100
- pyhygese 1s 98/100
- rustvrp 1s 91/100
- solverforge 1s 84/100
- solverforge-py 1s 58/100
- timefold 1s 14/100
10-second budget
- pyvrp 10s 100/100
- ortools 10s 100/100
- rustvrp 10s 100/100
- solverforge 10s 100/100
- timefold 10s 100/100
- vroom 10s 99/100
- pyhygese 10s 99/100
- solverforge-py 10s 58/100
60-second budget
- pyvrp 60s 100/100
- ortools 60s 100/100
- vroom 60s 100/100
- pyhygese 60s 100/100
- rustvrp 60s 100/100
- solverforge 60s 100/100
- timefold 60s 100/100
- solverforge-py 60s 76/100
First feasible gate
Instances first observed feasible at this tested budget, out of 100. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: pyvrp 0; vroom 0; ortools 0; pyhygese 0; rustvrp 0; solverforge 0; solverforge-py 24; timefold 0.
1-second budget
- pyvrp 1s 100/100
- vroom 1s 100/100
- ortools 1s 100/100
- pyhygese 1s 98/100
- rustvrp 1s 91/100
- solverforge 1s 84/100
- solverforge-py 1s 58/100
- timefold 1s 14/100
10-second budget
- pyvrp 10s 0/100
- ortools 10s 0/100
- rustvrp 10s 9/100
- solverforge 10s 16/100
- timefold 10s 86/100
- vroom 10s 0/100
- pyhygese 10s 1/100
- solverforge-py 10s 1/100
60-second budget
- pyvrp 60s 0/100
- ortools 60s 0/100
- vroom 60s 0/100
- pyhygese 60s 1/100
- rustvrp 60s 0/100
- solverforge 60s 0/100
- timefold 60s 0/100
- solverforge-py 60s 17/100
Mean gap to reference
Feasible results with a reference cost only. 0% matches the reference; lower is better. Rows follow feasibility-first ranking, not gap alone. Coverage and sample counts are shown on every row; means over different solved subsets are not like-for-like. Bars are scaled to this problem's largest mean gap, 81.91%.
1-second budget
- pyvrp 1s 4.16% 100/100 feasible · 100 gap samples
- vroom 1s 7.34% 100/100 feasible · 100 gap samples
- ortools 1s 10.43% 100/100 feasible · 100 gap samples
- pyhygese 1s 2.55% 98/100 feasible · 98 gap samples
- rustvrp 1s 12.94% 91/100 feasible · 91 gap samples
- solverforge 1s 5.53% 84/100 feasible · 84 gap samples
- solverforge-py 1s 6.02% 58/100 feasible · 58 gap samples
- timefold 1s 81.91% 14/100 feasible · 14 gap samples
10-second budget
- pyvrp 10s 1.43% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100
- ortools 10s 6.89% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100
- rustvrp 10s 5.83% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 91/100
- solverforge 10s 5.09% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 84/100
- timefold 10s 58.27% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 14/100
- vroom 10s 3.56% 99/100 feasible · 99 gap samples Earlier feasibility: 1s 100/100
- pyhygese 10s 1.59% 99/100 feasible · 99 gap samples Earlier feasibility: 1s 98/100
- solverforge-py 10s 5.86% 58/100 feasible · 58 gap samples Earlier feasibility: 1s 58/100
60-second budget
- pyvrp 60s 0.52% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 100/100
- ortools 60s 5.52% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 100/100
- vroom 60s 2.18% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 99/100
- pyhygese 60s 1.04% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 98/100 · 10s 99/100
- rustvrp 60s 3.69% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 91/100 · 10s 100/100
- solverforge 60s 4.95% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 84/100 · 10s 100/100
- timefold 60s 18.63% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 14/100 · 10s 100/100
- solverforge-py 60s 5.55% 76/100 feasible · 76 gap samples Earlier feasibility: 1s 58/100 · 10s 58/100
1-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| pyvrp 0.13.4 | 100 / 100 | 4.16% (100 samples) | 66008.62 (100 samples) | 1.59 s | 1.59 s | 93 / 100 |
| vroom 1.15.0 | 100 / 100 | 7.34% (100 samples) | 67335.07 (100 samples) | 1.64 s | 1.64 s | 88 / 100 |
| ortools 9.15.6755 | 100 / 100 | 10.43% (100 samples) | 69796.67 (100 samples) | 1.11 s | 1.11 s | 33 / 100 |
| pyhygese 0.1.0 | 98 / 100 | 2.55% (98 samples) | 63163.13 (98 samples) | 1.03 s | 1.03 s | 4 / 100 |
| rustvrp 1.25.0 | 91 / 100 | 12.94% (91 samples) | 62759.86 (91 samples) | 1.08 s | 1.10 s | 32 / 100 |
| solverforge 0.19.4 | 84 / 100 | 5.53% (84 samples) | 51220.73 (84 samples) | 1.03 s | 1.04 s | 4 / 100 |
| solverforge-py 0.6.6 | 58 / 100 | 6.02% (58 samples) | 40186.14 (58 samples) | 2.56 s | 4.05 s | 100 / 100 |
| timefold 2.1.0 | 14 / 100 | 81.91% (14 samples) | 37046.79 (14 samples) | 2.30 s | 2.22 s | 100 / 100 |
10-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| pyvrp 0.13.4 | 100 / 100 | 1.43% (100 samples) | 64209.58 (100 samples) | 10.57 s | 10.57 s | 20 / 100 |
| ortools 9.15.6755 | 100 / 100 | 6.89% (100 samples) | 66987.06 (100 samples) | 10.10 s | 10.10 s | 0 / 100 |
| rustvrp 1.25.0 | 100 / 100 | 5.83% (100 samples) | 66281.51 (100 samples) | 10.10 s | 10.10 s | 0 / 100 |
| solverforge 0.19.4 | 100 / 100 | 5.09% (100 samples) | 65958.22 (100 samples) | 10.04 s | 10.04 s | 0 / 100 |
| timefold 2.1.0 | 100 / 100 | 58.27% (100 samples) | 98201.94 (100 samples) | 11.18 s | 11.18 s | 88 / 100 |
| vroom 1.15.0 | 99 / 100 | 3.56% (99 samples) | 65383.10 (99 samples) | 10.37 s | 10.42 s | 10 / 100 |
| pyhygese 0.1.0 | 99 / 100 | 1.59% (99 samples) | 63587.58 (99 samples) | 10.09 s | 10.09 s | 0 / 100 |
| solverforge-py 0.6.6 | 58 / 100 | 5.86% (58 samples) | 41542.41 (58 samples) | 11.57 s | 13.09 s | 77 / 100 |
60-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| pyvrp 0.13.4 | 100 / 100 | 0.52% (100 samples) | 63507.71 (100 samples) | 60.59 s | 60.59 s | 0 / 100 |
| ortools 9.15.6755 | 100 / 100 | 5.52% (100 samples) | 66428.86 (100 samples) | 60.11 s | 60.11 s | 0 / 100 |
| vroom 1.15.0 | 100 / 100 | 2.18% (100 samples) | 64602.77 (100 samples) | 45.59 s | 45.59 s | 0 / 100 |
| pyhygese 0.1.0 | 100 / 100 | 1.04% (100 samples) | 63970.41 (100 samples) | 60.43 s | 60.43 s | 0 / 100 |
| rustvrp 1.25.0 | 100 / 100 | 3.69% (100 samples) | 65077.30 (100 samples) | 60.10 s | 60.10 s | 0 / 100 |
| solverforge 0.19.4 | 100 / 100 | 4.95% (100 samples) | 65863.43 (100 samples) | 60.04 s | 60.04 s | 0 / 100 |
| timefold 2.1.0 | 100 / 100 | 18.63% (100 samples) | 75127.77 (100 samples) | 61.18 s | 61.18 s | 0 / 100 |
| solverforge-py 0.6.6 | 76 / 100 | 5.55% (76 samples) | 51904.12 (76 samples) | 63.28 s | 66.20 s | 39 / 100 |
Run provenance
Full canonical nightly candidate · completed 2026-10-01T10:34:31.304855+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.
Run 3614f02a-0336-4153-a575-f434697506fb
Harness commit 3c1710510037.
Employee scheduling
Assign nurses to shifts while respecting hard rules. Among feasible schedules, lower penalty cost is better.
INRC-II · 42 instances · 504 recorded results · Completed
Reference values cover 0 of 42 instances in this run — the catalog holds 37 published values for this problem, on history/week tuples this run does not select — so their mean gap is not shown. The values come from INRC-II official test dataset (mobiz.vives.be/inrc2): best known bounds, not proven optima.
Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.
Feasible results
Instances solved to a hard-feasible solution, out of 42.
1-second budget
- solverforge 1s 42/42
- solverforge-py 1s 42/42
- timefold 1s 4/42
- ortools 1s 0/42
10-second budget
- solverforge 10s 42/42
- solverforge-py 10s 42/42
- timefold 10s 42/42
- ortools 10s 0/42
60-second budget
- solverforge 60s 42/42
- solverforge-py 60s 42/42
- timefold 60s 42/42
- ortools 60s 0/42
First feasible gate
Instances first observed feasible at this tested budget, out of 42. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: solverforge 0; solverforge-py 0; timefold 0; ortools 42.
1-second budget
- solverforge 1s 42/42
- solverforge-py 1s 42/42
- timefold 1s 4/42
- ortools 1s 0/42
10-second budget
- solverforge 10s 0/42
- solverforge-py 10s 0/42
- timefold 10s 38/42
- ortools 10s 0/42
60-second budget
- solverforge 60s 0/42
- solverforge-py 60s 0/42
- timefold 60s 0/42
- ortools 60s 0/42
Mean gap to reference
This run's instances are not in the published reference set, so no solver has a mean gap to show. Validated mean costs are in the tables below; the published values cover other instances of this problem, not these.
1-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| solverforge 0.19.4 | 42 / 42 | — (0 samples) | 71476.79 (42 samples) | 1.06 s | 1.06 s | 0 / 42 |
| solverforge-py 0.6.6 | 42 / 42 | — (0 samples) | 77952.02 (42 samples) | 2.47 s | 2.47 s | 42 / 42 |
| timefold 2.1.0 | 4 / 42 | — (0 samples) | 8032.50 (4 samples) | 2.47 s | 2.62 s | 42 / 42 |
| ortools 9.15.6755 | 0 / 42 | — (0 samples) | — (0 samples) | — | 1.32 s | 39 / 42 |
10-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| solverforge 0.19.4 | 42 / 42 | — (0 samples) | 50262.38 (42 samples) | 10.06 s | 10.06 s | 0 / 42 |
| solverforge-py 0.6.6 | 42 / 42 | — (0 samples) | 64976.55 (42 samples) | 11.22 s | 11.22 s | 21 / 42 |
| timefold 2.1.0 | 42 / 42 | — (0 samples) | 20580.95 (42 samples) | 11.58 s | 11.58 s | 42 / 42 |
| ortools 9.15.6755 | 0 / 42 | — (0 samples) | — (0 samples) | — | 10.24 s | 0 / 42 |
60-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| solverforge 0.19.4 | 42 / 42 | — (0 samples) | 19891.19 (42 samples) | 60.10 s | 60.10 s | 0 / 42 |
| solverforge-py 0.6.6 | 42 / 42 | — (0 samples) | 42119.29 (42 samples) | 61.25 s | 61.25 s | 0 / 42 |
| timefold 2.1.0 | 42 / 42 | — (0 samples) | 20029.64 (42 samples) | 61.58 s | 61.58 s | 0 / 42 |
| ortools 9.15.6755 | 0 / 42 | — (0 samples) | — (0 samples) | — | 60.30 s | 0 / 42 |
Run provenance
Full canonical nightly candidate · completed 2026-09-30T21:59:48.901833+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.
Run aeabc780-28fc-4e7d-b0fc-41fbbd2dbe97
Harness commit 3c1710510037.
Job-shop scheduling
Sequence jobs on machines while respecting operation order and machine capacity. Lower makespan is better.
JSPLIB · 162 instances · 1944 recorded results · Completed
Reference values for all 162 instances, from ScheduleOpt/benchmarks: proven optima where the source closed the instance, best known bounds otherwise.
Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.
Feasible results
Instances solved to a hard-feasible solution, out of 162.
1-second budget
- ortools 1s 162/162
- solverforge 1s 135/162
- solverforge-py 1s 114/162
- timefold 1s 104/162
10-second budget
- ortools 10s 162/162
- solverforge 10s 162/162
- solverforge-py 10s 162/162
- timefold 10s 155/162
60-second budget
- ortools 60s 162/162
- solverforge 60s 162/162
- solverforge-py 60s 162/162
- timefold 60s 162/162
First feasible gate
Instances first observed feasible at this tested budget, out of 162. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: ortools 0; solverforge 0; solverforge-py 0; timefold 0.
1-second budget
- ortools 1s 162/162
- solverforge 1s 135/162
- solverforge-py 1s 114/162
- timefold 1s 104/162
10-second budget
- ortools 10s 0/162
- solverforge 10s 27/162
- solverforge-py 10s 48/162
- timefold 10s 51/162
60-second budget
- ortools 60s 0/162
- solverforge 60s 0/162
- solverforge-py 60s 0/162
- timefold 60s 7/162
Mean gap to reference
Feasible results with a reference cost only. 0% matches the reference; lower is better. Rows follow feasibility-first ranking, not gap alone. Coverage and sample counts are shown on every row; means over different solved subsets are not like-for-like. Bars are scaled to this problem's largest mean gap, 30.82%.
1-second budget
- ortools 1s 12.09% 162/162 feasible · 162 gap samples
- solverforge 1s 15.31% 135/162 feasible · 135 gap samples
- solverforge-py 1s 13.85% 114/162 feasible · 114 gap samples
- timefold 1s 30.82% 104/162 feasible · 104 gap samples
10-second budget
- ortools 10s 8.12% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 162/162
- solverforge 10s 12.52% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 135/162
- solverforge-py 10s 12.87% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 114/162
- timefold 10s 25.43% 155/162 feasible · 155 gap samples Earlier feasibility: 1s 104/162
60-second budget
- ortools 60s 5.74% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 162/162 · 10s 162/162
- solverforge 60s 10.73% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 135/162 · 10s 162/162
- solverforge-py 60s 11.04% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 114/162 · 10s 162/162
- timefold 60s 13.76% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 104/162 · 10s 155/162
1-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| ortools 9.15.6755 | 162 / 162 | 12.09% (162 samples) | 2078.52 (162 samples) | 0.93 s | 0.93 s | 2 / 162 |
| solverforge 0.19.4 | 135 / 162 | 15.31% (135 samples) | 2260.84 (135 samples) | 1.07 s | 1.06 s | 19 / 162 |
| solverforge-py 0.6.6 | 114 / 162 | 13.85% (114 samples) | 2381.13 (114 samples) | 1.10 s | 1.08 s | 35 / 162 |
| timefold 2.1.0 | 104 / 162 | 30.82% (104 samples) | 1621.85 (104 samples) | 1.85 s | 1.88 s | 162 / 162 |
10-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| ortools 9.15.6755 | 162 / 162 | 8.12% (162 samples) | 2017.31 (162 samples) | 8.20 s | 8.20 s | 0 / 162 |
| solverforge 0.19.4 | 162 / 162 | 12.52% (162 samples) | 2086.43 (162 samples) | 10.06 s | 10.06 s | 2 / 162 |
| solverforge-py 0.6.6 | 162 / 162 | 12.87% (162 samples) | 2090.91 (162 samples) | 10.07 s | 10.07 s | 0 / 162 |
| timefold 2.1.0 | 155 / 162 | 25.43% (155 samples) | 2249.88 (155 samples) | 10.89 s | 10.89 s | 15 / 162 |
60-second budget
| Solver / tested version | Feasible | Mean gap | Mean cost | Runtime (feasible) | Mean runtime | Over budget |
|---|---|---|---|---|---|---|
| ortools 9.15.6755 | 162 / 162 | 5.74% (162 samples) | 1974.99 (162 samples) | 44.08 s | 44.08 s | 0 / 162 |
| solverforge 0.19.4 | 162 / 162 | 10.73% (162 samples) | 2057.89 (162 samples) | 60.04 s | 60.04 s | 0 / 162 |
| solverforge-py 0.6.6 | 162 / 162 | 11.04% (162 samples) | 2062.17 (162 samples) | 60.09 s | 60.09 s | 0 / 162 |
| timefold 2.1.0 | 162 / 162 | 13.76% (162 samples) | 2259.20 (162 samples) | 60.89 s | 60.89 s | 0 / 162 |
Run provenance
Full canonical nightly candidate · completed 2026-10-01T06:42:42.762173+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.
Run 2137acee-ec6d-4c85-a6ff-206a934f86c0
Harness commit 3c1710510037.
Evidence and methodology
Gates are independent runs with requested budgets, not samples of one continuous search. First feasible gate means the earliest tested budget that returned a hard-feasible result for that instance; it is not the instant feasibility was discovered. No first-incumbent timestamps were recorded.
Runtime is measured wall-clock duration, shown only as a diagnostic. “Over budget” counts results flagged by the harness’s wall-time tolerance; late returned solutions are retained, so these gates are not strict deadline guarantees. The tables start at 1s; expand the later gates to inspect improvement.
Download the recorded results and run metadata (JSON) · How we benchmark SolverForge · Benchmark source