Solver Benchmarks

sTiles against five sparse direct solvers across the full 60-matrix suite. All results shown here are from the sTiles paper.

Every matrix in the suite was factored by sTiles and five reference sparse direct solvers. For each matrix and solver we report the best factorization time over a sweep of thread counts. Times below are the same measurements reported in the paper.

60 matrices 6 solvers Intel 2×Xeon Gold · 40 cores best over 1–40 core sweep Cholesky factorization
Factorization time across all matrices
Best factorization time for every matrix and solver

Best factorization time (minimum over the swept core counts) for every matrix in the 60-matrix suite and every solver, on the Intel node. Matrices are sorted along the horizontal axis by sTiles' time, so the sTiles curve is monotonic; a competitor marker below it is faster than sTiles on that matrix, above it slower. PARDISO is the closest solver: it is faster than sTiles on 15 matrices, all below 0.5 s; sTiles is faster on the other 45, including all 12 above 0.5 s. MUMPS, PaStiX, CHOLMOD, and symPACK are slower than sTiles on 97 to 100% of the matrices each of them factors.

Per-matrix best factorization time (s) · Intel node
#matrixsTilesPARDISOMUMPSCHOLMODPaStiXsymPACK
1inla_graph_sem_n20000.00050.00150.00690.00150.00540.0095
2inla_graph_sem_n50000.00080.00350.0150.00410.0150.020
3inla_graph_8rtKSK0.00100.00110.00580.00090.0783.00
4nasa47040.00130.00120.00580.00690.00820.0097
5gyro_m0.00130.00130.0130.0150.0230.154
6bcsstk150.00160.00150.00960.0100.00780.013
7gyro_k0.00180.00230.0250.0310.0570.014
8inla_graph_sh7Pgi0.00260.00190.00860.00380.009323.0
9inla_graph_pedigree0.00290.00700.0820.00560.041—
10inla_graph_sem_n200000.00310.0130.0540.0210.0630.080
11inla_graph_diff0.00320.00510.0180.0480.0290.034
12msc108480.00360.00390.0220.0330.0710.017
13inla_graph_net8143810.00380.00630.0320.0130.0311.21
14inla_graph_ayaLRw0.00380.00590.0270.0140.0290.228
15msc230520.00470.00500.0320.0510.0720.021
16thermal10.00520.00680.0360.1110.0430.021
17inla_graph_lgm_10010_bw20.0100.0130.0360.0440.0480.023
18nasasrb0.0120.0140.0610.1720.1630.037
19oilpan0.0130.0120.0650.1600.2190.050
20inla_graph_sem_n1000000.0140.0680.2760.1620.3530.400
21thermomech_dM0.0150.0170.0900.2870.1250.029
22ct20stif0.0180.0200.0730.1520.1910.069
23s3dkt3m20.0190.0210.1040.2310.2480.103
24smt0.0210.0220.0920.1400.2530.095
25inla_graph_ferris0.0210.0130.0360.1430.1160.087
26s3dkq4m20.0220.0210.1110.2470.2910.101
27inla_graph_animal10.0230.1810.1330.1100.3910.613
28apache10.0240.0230.0810.2730.0790.127
29inla_graph_bern_spd0.0290.0440.1260.5110.3060.775
30bmw7st_10.0300.0330.1260.3940.4600.175
31inla_graph_83o4NNNo0.0310.00820.0530.1200.0960.023
32inla_graph_stcov0.0380.1260.8271.031.0001.39
33inla_graph_spacetime0.0470.0890.1330.1880.1430.208
34nd3k0.0510.0810.1810.1650.2810.375
35pwtk0.0520.0550.2600.6720.6730.139
36inla_graph_lidense0.0560.1110.3770.4961.551.15
37crankseg_10.0620.0570.2470.3950.6950.158
38bmw3_20.0740.0700.2320.6850.7320.320
39boneS010.0840.0830.2270.5330.5460.552
40crankseg_20.0860.0810.2950.5240.9210.230
41tmt_sym0.0870.0890.2541.200.4530.176
42ecology20.1070.1280.3261.530.4850.359
43af_shell30.1600.1520.4321.371.190.409
44consph0.1880.2060.5120.7600.7191.81
45nd6k0.2080.3120.4750.4860.6792.86
46inline_10.2970.3400.8262.312.540.481
47inla_graph_yU0G1u0.2980.2840.7121.55311—
48inla_graph_net16287600.3620.2290.9853.703.520.819
49boneS100.6080.6291.273.693.921.13
50inla_graph_lgm_48600_bw20.7239.434.342.368.109.45
51nd12k0.8911.471.251.481.839.27
52inla_graph_lgm_100200_bw11.132.481.982.224.651.85
53inla_graph_lgm_50000_bw150001.445.322.595.302.8128.9
54inla_graph_lgm_100200_bw22.387.634.583.546.738.34
55bone0106.447.887.0311.89.5621.8
56audikw_16.6712.29.1814.710.547.3
57Fault_6397.8817.110.112.410.9110
58inla_graph_lgm_50400_bw211.937.133.57.1614.976.7
59Emilia_92313.324.815.719.919.8115
60inla_graph_animal216.620935.725.072.1432
fastest solver on that matrix · — marks a matrix a solver does not factor

Best factorization time in seconds on the Intel node (minimum over the swept core counts) for every matrix and solver, sorted by sTiles' time to match the figure above. The fastest solver on each matrix is highlighted.

Speedup of sTiles over PARDISO
Per-matrix speedup of sTiles over PARDISO

Per-matrix speedup of sTiles over PARDISO against matrix cost (both axes logarithmic; a value above the 1× line means sTiles is faster). sTiles is faster on 45 matrices and slower on 15, with a geometric-mean ratio of 1.50×. The two outcomes are asymmetric: every loss is on an inexpensive matrix, all 15 under 0.15 s and the largest margin 0.133 s, while the wins are on the expensive matrices and reach minutes.

Parallel efficiency to 40 cores
Strong scaling of sTiles vs PARDISO

Parallel efficiency (speedup over one core divided by the core count; 1.0 is ideal) of a single factorization on three large finite-element matrices, sTiles (solid) against PARDISO (dashed). sTiles holds above 0.85 through 16 cores, where PARDISO has already fallen to 0.5–0.6; both drop sharply past 32 cores (shaded band), the socket's memory-bandwidth ceiling for sparse Cholesky.