Measured performance

Benchmarks

Latest full benchmark results from solverforge-bench, grouped by problem and labeled with the versions actually tested.

Latest completed full canonical nightly runs, with the versions actually tested. Quick and partial runs are excluded; these are not measurements of newer releases.

Read the gates from left to right: 1s → 10s → 60s. First, compare how many instances reach a valid solution at the earliest tested gate. Then compare quality at that same gate; later gates show improvement. A short process runtime is not a win.

Each gate ranks feasibility → earlier-gate feasibility → solution quality, never process duration. Among solvers with equal current coverage, more valid solutions at 1s wins first, then at 10s where applicable; quality breaks the remaining tie. 0% gap matches the reference. Without references, lower validated mean cost is shown instead. Quality averages cover feasible results only; different solved subsets are not a like-for-like quality comparison. Problems are compared separately.

Capacitated vehicle routing

Assign customers to vehicle routes while respecting capacity. Lower route cost is better.

CVRPLIB-X · 100 instances · 2400 recorded results · Completed

Reference values for all 100 instances, from CVRPLIB (Uchoa et al. 2017 Set X): proven optima where the source closed the instance, best known bounds otherwise.

Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.

1-second budget 10-second budget 60-second budget

Feasible results

Instances solved to a hard-feasible solution, out of 100.

1-second budget

  1. pyvrp 1s 100/100
  2. vroom 1s 100/100
  3. ortools 1s 100/100
  4. pyhygese 1s 98/100
  5. rustvrp 1s 91/100
  6. solverforge 1s 84/100
  7. solverforge-py 1s 58/100
  8. timefold 1s 14/100

10-second budget

  1. pyvrp 10s 100/100
  2. ortools 10s 100/100
  3. rustvrp 10s 100/100
  4. solverforge 10s 100/100
  5. timefold 10s 100/100
  6. vroom 10s 99/100
  7. pyhygese 10s 99/100
  8. solverforge-py 10s 58/100

60-second budget

  1. pyvrp 60s 100/100
  2. ortools 60s 100/100
  3. vroom 60s 100/100
  4. pyhygese 60s 100/100
  5. rustvrp 60s 100/100
  6. solverforge 60s 100/100
  7. timefold 60s 100/100
  8. solverforge-py 60s 76/100

First feasible gate

Instances first observed feasible at this tested budget, out of 100. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: pyvrp 0; vroom 0; ortools 0; pyhygese 0; rustvrp 0; solverforge 0; solverforge-py 24; timefold 0.

1-second budget

  1. pyvrp 1s 100/100
  2. vroom 1s 100/100
  3. ortools 1s 100/100
  4. pyhygese 1s 98/100
  5. rustvrp 1s 91/100
  6. solverforge 1s 84/100
  7. solverforge-py 1s 58/100
  8. timefold 1s 14/100

10-second budget

  1. pyvrp 10s 0/100
  2. ortools 10s 0/100
  3. rustvrp 10s 9/100
  4. solverforge 10s 16/100
  5. timefold 10s 86/100
  6. vroom 10s 0/100
  7. pyhygese 10s 1/100
  8. solverforge-py 10s 1/100

60-second budget

  1. pyvrp 60s 0/100
  2. ortools 60s 0/100
  3. vroom 60s 0/100
  4. pyhygese 60s 1/100
  5. rustvrp 60s 0/100
  6. solverforge 60s 0/100
  7. timefold 60s 0/100
  8. solverforge-py 60s 17/100

Mean gap to reference

Feasible results with a reference cost only. 0% matches the reference; lower is better. Rows follow feasibility-first ranking, not gap alone. Coverage and sample counts are shown on every row; means over different solved subsets are not like-for-like. Bars are scaled to this problem's largest mean gap, 81.91%.

1-second budget

  1. pyvrp 1s 4.16% 100/100 feasible · 100 gap samples
  2. vroom 1s 7.34% 100/100 feasible · 100 gap samples
  3. ortools 1s 10.43% 100/100 feasible · 100 gap samples
  4. pyhygese 1s 2.55% 98/100 feasible · 98 gap samples
  5. rustvrp 1s 12.94% 91/100 feasible · 91 gap samples
  6. solverforge 1s 5.53% 84/100 feasible · 84 gap samples
  7. solverforge-py 1s 6.02% 58/100 feasible · 58 gap samples
  8. timefold 1s 81.91% 14/100 feasible · 14 gap samples

10-second budget

  1. pyvrp 10s 1.43% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100
  2. ortools 10s 6.89% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100
  3. rustvrp 10s 5.83% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 91/100
  4. solverforge 10s 5.09% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 84/100
  5. timefold 10s 58.27% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 14/100
  6. vroom 10s 3.56% 99/100 feasible · 99 gap samples Earlier feasibility: 1s 100/100
  7. pyhygese 10s 1.59% 99/100 feasible · 99 gap samples Earlier feasibility: 1s 98/100
  8. solverforge-py 10s 5.86% 58/100 feasible · 58 gap samples Earlier feasibility: 1s 58/100

60-second budget

  1. pyvrp 60s 0.52% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 100/100
  2. ortools 60s 5.52% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 100/100
  3. vroom 60s 2.18% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 100/100 · 10s 99/100
  4. pyhygese 60s 1.04% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 98/100 · 10s 99/100
  5. rustvrp 60s 3.69% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 91/100 · 10s 100/100
  6. solverforge 60s 4.95% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 84/100 · 10s 100/100
  7. timefold 60s 18.63% 100/100 feasible · 100 gap samples Earlier feasibility: 1s 14/100 · 10s 100/100
  8. solverforge-py 60s 5.55% 76/100 feasible · 76 gap samples Earlier feasibility: 1s 58/100 · 10s 58/100

1-second budget

Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
pyvrp 0.13.4 100 / 100 4.16% (100 samples) 66008.62 (100 samples) 1.59 s 1.59 s 93 / 100
vroom 1.15.0 100 / 100 7.34% (100 samples) 67335.07 (100 samples) 1.64 s 1.64 s 88 / 100
ortools 9.15.6755 100 / 100 10.43% (100 samples) 69796.67 (100 samples) 1.11 s 1.11 s 33 / 100
pyhygese 0.1.0 98 / 100 2.55% (98 samples) 63163.13 (98 samples) 1.03 s 1.03 s 4 / 100
rustvrp 1.25.0 91 / 100 12.94% (91 samples) 62759.86 (91 samples) 1.08 s 1.10 s 32 / 100
solverforge 0.19.4 84 / 100 5.53% (84 samples) 51220.73 (84 samples) 1.03 s 1.04 s 4 / 100
solverforge-py 0.6.6 58 / 100 6.02% (58 samples) 40186.14 (58 samples) 2.56 s 4.05 s 100 / 100
timefold 2.1.0 14 / 100 81.91% (14 samples) 37046.79 (14 samples) 2.30 s 2.22 s 100 / 100
10-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
pyvrp 0.13.4 100 / 100 1.43% (100 samples) 64209.58 (100 samples) 10.57 s 10.57 s 20 / 100
ortools 9.15.6755 100 / 100 6.89% (100 samples) 66987.06 (100 samples) 10.10 s 10.10 s 0 / 100
rustvrp 1.25.0 100 / 100 5.83% (100 samples) 66281.51 (100 samples) 10.10 s 10.10 s 0 / 100
solverforge 0.19.4 100 / 100 5.09% (100 samples) 65958.22 (100 samples) 10.04 s 10.04 s 0 / 100
timefold 2.1.0 100 / 100 58.27% (100 samples) 98201.94 (100 samples) 11.18 s 11.18 s 88 / 100
vroom 1.15.0 99 / 100 3.56% (99 samples) 65383.10 (99 samples) 10.37 s 10.42 s 10 / 100
pyhygese 0.1.0 99 / 100 1.59% (99 samples) 63587.58 (99 samples) 10.09 s 10.09 s 0 / 100
solverforge-py 0.6.6 58 / 100 5.86% (58 samples) 41542.41 (58 samples) 11.57 s 13.09 s 77 / 100
60-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
pyvrp 0.13.4 100 / 100 0.52% (100 samples) 63507.71 (100 samples) 60.59 s 60.59 s 0 / 100
ortools 9.15.6755 100 / 100 5.52% (100 samples) 66428.86 (100 samples) 60.11 s 60.11 s 0 / 100
vroom 1.15.0 100 / 100 2.18% (100 samples) 64602.77 (100 samples) 45.59 s 45.59 s 0 / 100
pyhygese 0.1.0 100 / 100 1.04% (100 samples) 63970.41 (100 samples) 60.43 s 60.43 s 0 / 100
rustvrp 1.25.0 100 / 100 3.69% (100 samples) 65077.30 (100 samples) 60.10 s 60.10 s 0 / 100
solverforge 0.19.4 100 / 100 4.95% (100 samples) 65863.43 (100 samples) 60.04 s 60.04 s 0 / 100
timefold 2.1.0 100 / 100 18.63% (100 samples) 75127.77 (100 samples) 61.18 s 61.18 s 0 / 100
solverforge-py 0.6.6 76 / 100 5.55% (76 samples) 51904.12 (76 samples) 63.28 s 66.20 s 39 / 100
Run provenance

Full canonical nightly candidate · completed 2026-10-01T10:34:31.304855+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.

Run 3614f02a-0336-4153-a575-f434697506fb
Harness commit 3c1710510037.

Employee scheduling

Assign nurses to shifts while respecting hard rules. Among feasible schedules, lower penalty cost is better.

INRC-II · 42 instances · 504 recorded results · Completed

Reference values cover 0 of 42 instances in this run — the catalog holds 37 published values for this problem, on history/week tuples this run does not select — so their mean gap is not shown. The values come from INRC-II official test dataset (mobiz.vives.be/inrc2): best known bounds, not proven optima.

Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.

1-second budget 10-second budget 60-second budget

Feasible results

Instances solved to a hard-feasible solution, out of 42.

1-second budget

  1. solverforge 1s 42/42
  2. solverforge-py 1s 42/42
  3. timefold 1s 4/42
  4. ortools 1s 0/42

10-second budget

  1. solverforge 10s 42/42
  2. solverforge-py 10s 42/42
  3. timefold 10s 42/42
  4. ortools 10s 0/42

60-second budget

  1. solverforge 60s 42/42
  2. solverforge-py 60s 42/42
  3. timefold 60s 42/42
  4. ortools 60s 0/42

First feasible gate

Instances first observed feasible at this tested budget, out of 42. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: solverforge 0; solverforge-py 0; timefold 0; ortools 42.

1-second budget

  1. solverforge 1s 42/42
  2. solverforge-py 1s 42/42
  3. timefold 1s 4/42
  4. ortools 1s 0/42

10-second budget

  1. solverforge 10s 0/42
  2. solverforge-py 10s 0/42
  3. timefold 10s 38/42
  4. ortools 10s 0/42

60-second budget

  1. solverforge 60s 0/42
  2. solverforge-py 60s 0/42
  3. timefold 60s 0/42
  4. ortools 60s 0/42

Mean gap to reference

This run's instances are not in the published reference set, so no solver has a mean gap to show. Validated mean costs are in the tables below; the published values cover other instances of this problem, not these.

1-second budget

Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
solverforge 0.19.4 42 / 42 — (0 samples) 71476.79 (42 samples) 1.06 s 1.06 s 0 / 42
solverforge-py 0.6.6 42 / 42 — (0 samples) 77952.02 (42 samples) 2.47 s 2.47 s 42 / 42
timefold 2.1.0 4 / 42 — (0 samples) 8032.50 (4 samples) 2.47 s 2.62 s 42 / 42
ortools 9.15.6755 0 / 42 — (0 samples) — (0 samples) — 1.32 s 39 / 42
10-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
solverforge 0.19.4 42 / 42 — (0 samples) 50262.38 (42 samples) 10.06 s 10.06 s 0 / 42
solverforge-py 0.6.6 42 / 42 — (0 samples) 64976.55 (42 samples) 11.22 s 11.22 s 21 / 42
timefold 2.1.0 42 / 42 — (0 samples) 20580.95 (42 samples) 11.58 s 11.58 s 42 / 42
ortools 9.15.6755 0 / 42 — (0 samples) — (0 samples) — 10.24 s 0 / 42
60-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
solverforge 0.19.4 42 / 42 — (0 samples) 19891.19 (42 samples) 60.10 s 60.10 s 0 / 42
solverforge-py 0.6.6 42 / 42 — (0 samples) 42119.29 (42 samples) 61.25 s 61.25 s 0 / 42
timefold 2.1.0 42 / 42 — (0 samples) 20029.64 (42 samples) 61.58 s 61.58 s 0 / 42
ortools 9.15.6755 0 / 42 — (0 samples) — (0 samples) — 60.30 s 0 / 42
Run provenance

Full canonical nightly candidate · completed 2026-09-30T21:59:48.901833+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.

Run aeabc780-28fc-4e7d-b0fc-41fbbd2dbe97
Harness commit 3c1710510037.

Job-shop scheduling

Sequence jobs on machines while respecting operation order and machine capacity. Lower makespan is better.

JSPLIB · 162 instances · 1944 recorded results · Completed

Reference values for all 162 instances, from ScheduleOpt/benchmarks: proven optima where the source closed the instance, best known bounds otherwise.

Earliest feasible gate first; compare feasibility and quality within each gate. Runtime does not rank solutions.

1-second budget 10-second budget 60-second budget

Feasible results

Instances solved to a hard-feasible solution, out of 162.

1-second budget

  1. ortools 1s 162/162
  2. solverforge 1s 135/162
  3. solverforge-py 1s 114/162
  4. timefold 1s 104/162

10-second budget

  1. ortools 10s 162/162
  2. solverforge 10s 162/162
  3. solverforge-py 10s 162/162
  4. timefold 10s 155/162

60-second budget

  1. ortools 60s 162/162
  2. solverforge 60s 162/162
  3. solverforge-py 60s 162/162
  4. timefold 60s 162/162

First feasible gate

Instances first observed feasible at this tested budget, out of 162. Each instance counts once: earlier is better. Independent budget runs, not first-solution timestamps; late returns are retained. Never feasible: ortools 0; solverforge 0; solverforge-py 0; timefold 0.

1-second budget

  1. ortools 1s 162/162
  2. solverforge 1s 135/162
  3. solverforge-py 1s 114/162
  4. timefold 1s 104/162

10-second budget

  1. ortools 10s 0/162
  2. solverforge 10s 27/162
  3. solverforge-py 10s 48/162
  4. timefold 10s 51/162

60-second budget

  1. ortools 60s 0/162
  2. solverforge 60s 0/162
  3. solverforge-py 60s 0/162
  4. timefold 60s 7/162

Mean gap to reference

Feasible results with a reference cost only. 0% matches the reference; lower is better. Rows follow feasibility-first ranking, not gap alone. Coverage and sample counts are shown on every row; means over different solved subsets are not like-for-like. Bars are scaled to this problem's largest mean gap, 30.82%.

1-second budget

  1. ortools 1s 12.09% 162/162 feasible · 162 gap samples
  2. solverforge 1s 15.31% 135/162 feasible · 135 gap samples
  3. solverforge-py 1s 13.85% 114/162 feasible · 114 gap samples
  4. timefold 1s 30.82% 104/162 feasible · 104 gap samples

10-second budget

  1. ortools 10s 8.12% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 162/162
  2. solverforge 10s 12.52% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 135/162
  3. solverforge-py 10s 12.87% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 114/162
  4. timefold 10s 25.43% 155/162 feasible · 155 gap samples Earlier feasibility: 1s 104/162

60-second budget

  1. ortools 60s 5.74% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 162/162 · 10s 162/162
  2. solverforge 60s 10.73% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 135/162 · 10s 162/162
  3. solverforge-py 60s 11.04% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 114/162 · 10s 162/162
  4. timefold 60s 13.76% 162/162 feasible · 162 gap samples Earlier feasibility: 1s 104/162 · 10s 155/162

1-second budget

Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
ortools 9.15.6755 162 / 162 12.09% (162 samples) 2078.52 (162 samples) 0.93 s 0.93 s 2 / 162
solverforge 0.19.4 135 / 162 15.31% (135 samples) 2260.84 (135 samples) 1.07 s 1.06 s 19 / 162
solverforge-py 0.6.6 114 / 162 13.85% (114 samples) 2381.13 (114 samples) 1.10 s 1.08 s 35 / 162
timefold 2.1.0 104 / 162 30.82% (104 samples) 1621.85 (104 samples) 1.85 s 1.88 s 162 / 162
10-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
ortools 9.15.6755 162 / 162 8.12% (162 samples) 2017.31 (162 samples) 8.20 s 8.20 s 0 / 162
solverforge 0.19.4 162 / 162 12.52% (162 samples) 2086.43 (162 samples) 10.06 s 10.06 s 2 / 162
solverforge-py 0.6.6 162 / 162 12.87% (162 samples) 2090.91 (162 samples) 10.07 s 10.07 s 0 / 162
timefold 2.1.0 155 / 162 25.43% (155 samples) 2249.88 (155 samples) 10.89 s 10.89 s 15 / 162
60-second budget
Solver / tested versionFeasibleMean gapMean costRuntime (feasible)Mean runtimeOver budget
ortools 9.15.6755 162 / 162 5.74% (162 samples) 1974.99 (162 samples) 44.08 s 44.08 s 0 / 162
solverforge 0.19.4 162 / 162 10.73% (162 samples) 2057.89 (162 samples) 60.04 s 60.04 s 0 / 162
solverforge-py 0.6.6 162 / 162 11.04% (162 samples) 2062.17 (162 samples) 60.09 s 60.09 s 0 / 162
timefold 2.1.0 162 / 162 13.76% (162 samples) 2259.20 (162 samples) 60.89 s 60.89 s 0 / 162
Run provenance

Full canonical nightly candidate · completed 2026-10-01T06:42:42.762173+00:00 (UTC). The warehouse's publication checks passed; the benchmark repository was clean.

Run 2137acee-ec6d-4c85-a6ff-206a934f86c0
Harness commit 3c1710510037.

Evidence and methodology

Gates are independent runs with requested budgets, not samples of one continuous search. First feasible gate means the earliest tested budget that returned a hard-feasible result for that instance; it is not the instant feasibility was discovered. No first-incumbent timestamps were recorded.

Runtime is measured wall-clock duration, shown only as a diagnostic. “Over budget” counts results flagged by the harness’s wall-time tolerance; late returned solutions are retained, so these gates are not strict deadline guarantees. The tables start at 1s; expand the later gates to inspect improvement.

Download the recorded results and run metadata (JSON) · How we benchmark SolverForge · Benchmark source