Solver Benchmarks

sTiles against five sparse direct solvers across the full 88-matrix suite. All results shown here are from the sTiles paper.

Every matrix in the suite was factored by sTiles and five reference sparse direct solvers. For each matrix and solver we report the best factorization time over a sweep of thread counts. Times below are the same measurements reported in the paper.

88 matrices 6 solvers Intel 2×Xeon Gold · 40 cores best over 1–40 core sweep Cholesky factorization
Factorization time across all matrices
Best factorization time for every matrix and solver

Best factorization time (minimum over the swept core counts) for every matrix in the 88-matrix suite and every solver, on the Intel node. Matrices are sorted along the horizontal axis by sTiles' time, so the sTiles curve is monotonic; a competitor marker below it is faster than sTiles on that matrix, above it slower. PARDISO tracks sTiles closely, ahead on the inexpensive matrices and behind on the expensive ones; MUMPS, PaStiX, CHOLMOD, and symPACK are slower across most of the suite.

Per-matrix best factorization time (s) · Intel node
#matrixsTilesPARDISOMUMPSCHOLMODPaStiXsymPACK
1bcsstk100.00020.00020.00120.00070.00150.0043
21138_bus0.00040.00010.00080.00010.00050.020
3inla_graph_sem_n20000.00050.00150.00690.00150.00540.0095
4bcsstk110.00080.00030.00200.00160.00250.0059
5inla_graph_sem_n50000.00080.00350.0150.00410.0150.020
6bcsstk090.00090.00030.00190.00170.00120.0056
7inla_graph_8rtKSK0.00100.00110.00580.00090.0783.00
8gyro_m0.00130.00120.0130.0150.0230.154
9bcsstk080.00140.00040.00160.00200.00110.0076
10bcsstk140.00140.00050.00340.00280.00380.0053
11nasa29100.00140.00090.00540.00470.0100.0076
12nasa47040.00240.00120.00580.00690.00820.0097
13inla_graph_sh7Pgi0.00240.00200.00860.00380.009323.0
14gyro_k0.00270.00220.0250.0310.0570.014
15inla_graph_pedigree0.00290.00700.0820.00560.041
16inla_graph_sem_n200000.00310.0130.0540.0210.0630.080
17bcsstk150.00320.00150.00960.0100.00780.013
18inla_graph_diff0.00320.00510.0180.0480.0290.034
19bundle10.00330.00280.0290.0230.0590.051
20msc108480.00360.00390.0220.0330.0710.017
21bcsstk180.00360.00230.0100.0210.0120.021
22inla_graph_net8143810.00380.00630.0320.0130.0311.21
23inla_graph_ayaLRw0.00380.00590.0270.0140.0290.228
24bcsstk170.00390.00240.0150.0200.0260.013
25bcsstk130.00420.00150.00680.00490.00640.010
26msc230520.00470.00500.0320.0510.0720.021
27bcsstk160.00480.00160.0170.0120.0180.012
28thermal10.00520.00680.0360.1110.0430.021
29thermomech_TK0.0110.00920.0560.1400.0600.021
30thermomech_TC0.0110.00930.0600.1410.0610.021
31nasasrb0.0120.0140.0610.1720.1630.037
32oilpan0.0130.0120.0650.1600.2190.050
33inla_graph_sem_n1000000.0140.0680.2760.1620.3530.400
34thermomech_dM0.0150.0170.0900.2870.1250.029
35inla_graph_lgm_10010_bw20.0150.0120.0360.0440.0480.023
36ct20stif0.0180.0200.0730.1520.1910.069
37inla_graph_ferris0.0210.0120.0360.1430.1160.087
38s3dkt3m20.0210.0210.1040.2310.2480.103
39s3dkq4m20.0220.0210.1110.2470.2910.101
40inla_graph_animal10.0230.1810.1330.1100.3910.613
41apache10.0240.0230.0810.2730.0790.127
42smt0.0270.0210.0920.1400.2530.095
43G2_circuit0.0290.0190.0660.2760.0800.098
44inla_graph_bern_spd0.0290.0440.1260.5110.3060.775
45bmw7st_10.0350.0340.1260.3940.4600.175
46inla_graph_83o4NNNo0.0370.00810.0530.1200.0960.023
47inla_graph_stcov0.0380.1260.8271.031.0001.39
48hood0.0460.0300.1680.4900.6820.099
49inla_graph_spacetime0.0470.0890.1330.1880.1430.208
50inla_graph_lidense0.0560.1110.3770.4961.551.15
51crankseg_10.0600.0570.2470.3950.6950.158
52m_t10.0620.0470.1920.4210.5920.150
53pwtk0.0650.0550.2600.6720.6730.139
54thread0.0730.0570.1920.2720.3810.425
55parabolic_fem0.0750.0580.2090.9080.3920.107
56bmw3_20.0830.0740.2320.6850.7320.320
57tmt_sym0.0870.0890.2541.200.4530.176
58ship_0010.0880.0220.1100.1830.2750.071
59boneS010.0920.0820.2270.5330.5460.552
60crankseg_20.0980.0800.2950.5240.9210.230
61shipsec10.1070.0720.2140.4940.5651.25
62ecology20.1070.1280.3261.530.4850.359
63shipsec80.1140.0780.2020.4370.4960.551
64nd3k0.1290.0800.1810.1650.2810.375
65shipsec50.1730.1150.3090.7590.7140.488
66af_shell30.1810.1470.4321.371.190.409
67ship_0030.1930.1470.3660.6630.6160.809
68consph0.2440.2030.5120.7600.7191.81
69ldoor0.2450.2020.7292.433.081.09
70af_0_k1010.2900.1410.4671.421.210.412
71inla_graph_yU0G1u0.2950.2860.7121.55311
72inline_10.2970.3400.8262.312.540.481
73nd6k0.3120.3010.4750.4860.6792.86
74inla_graph_net16287600.3870.2200.9853.703.520.819
75apache20.6450.4520.8012.690.9371.71
76inla_graph_lgm_48600_bw20.7239.434.342.368.109.45
77boneS100.7620.6201.273.693.921.13
78offshore0.7840.2710.4521.140.5290.653
79nd12k0.8911.471.251.481.839.27
80inla_graph_lgm_100200_bw11.132.481.982.224.651.85
81inla_graph_lgm_50000_bw150001.445.322.595.302.8128.9
82inla_graph_lgm_100200_bw22.387.634.583.546.738.34
83bone0106.447.887.0311.89.5621.8
84audikw_16.6712.29.1814.710.547.3
85Fault_6397.8817.110.112.410.9110
86inla_graph_lgm_50400_bw211.937.133.57.1614.976.7
87Emilia_92313.324.815.719.919.8115
88inla_graph_animal216.620935.725.072.1432
fastest solver on that matrix · — marks a matrix a solver does not factor

Best factorization time in seconds on the Intel node (minimum over the swept core counts) for every matrix and solver, sorted by sTiles' time to match the figure above. The fastest solver on each matrix is highlighted.

Speedup of sTiles over PARDISO
Per-matrix speedup of sTiles over PARDISO

Per-matrix speedup of sTiles over PARDISO against matrix cost (both axes logarithmic; a value above the 1× line means sTiles is faster). By count the contest is a near-tie: sTiles is slower on 54 matrices and faster on 34, yet the geometric-mean ratio (1.03×) already favors sTiles, because its fewer wins are the larger ones. Every loss is on a cheap matrix (all but one under 0.5 s), while the wins are on the expensive matrices and reach minutes.

Strong scaling to 40 cores
Strong scaling of sTiles vs PARDISO

Strong scaling (speedup over one core) of a single factorization on three large finite-element matrices, sTiles (solid) against PARDISO (dashed), with the ideal line. Both flatten near 32 cores (shaded band), the socket's memory-bandwidth ceiling for sparse Cholesky.