← all workloads

autodiff

n differentiable functions arranged in bounded-depth groups, each group differentiated in both forward and reverse mode, plus a differentiable generic. Stresses the autodiff IR transform (inside linkAndOptimizeIR) and the front-end checking of [Differentiable]. Scales by breadth. Scaling null: n scales group breadth with depth bounded by _GROUP_DEPTH, so ideal autodiff cost is O(n).

bucket: autodiff  ·  mode: target  ·  flags: -target spirv -emit-spirv-directly

Phase composition vs N (stacked sub-counters)

compileInner split into phase buckets (named leaves + (self) residuals) stacked across the sweep sizes — the top edge is compileInner, so you can see which phase drives the scaling.

autodiff — phase composition vs N (v2026.13, median ms) autodiff 7.6× over N 25→200 0.0 557 1114 25 50 100 200 N autodiff — parseTranslationUnit autodiff — SemanticChecking autodiff — generateIR autodiff — frontEndExecute (self) autodiff — specializeModule autodiff — simplifyIR autodiff — linkIR autodiff — unrollLoopsInModule autodiff — legalizeResourceTypes autodiff — legalizeExistentialTypeLayout autodiff — performMandatoryEarlyInlining autodiff — performForceInlining autodiff — linkAndOptimizeIR (self) autodiff — generateOutput (self) autodiff — compileInner (self) phase buckets parseTranslationUnit SemanticChecking generateIR frontEndExecute (self) specializeModule simplifyIR linkIR unrollLoopsInModule legalizeResourceTypes legalizeExistentialTypeLayout performMandatoryEarlyInlining performForceInlining linkAndOptimizeIR (self) emitEntryPointsSourceFromIR generateOutput (self) compileInner (self)

Scaling analysis

floor-subtracted power-law fit (t − floor) = a·Nk; floor = the minimal workload (fixed per-compile cost), k the global exponent, top-2× the local high-end doubling ratio.

N rangefloor (ms)k (work)fit R²t(Nmin)t(Nmax)top-2×
25–200101.010.99413610322.26×

Growth attribution (N=25 → N=200)

compileInner grows by 896 ms across the sweep; the mutually-exclusive phase buckets below partition that growth exactly (no nested-timer double counting). × lin is the same metric as the top-level panels, per bucket: the end point vs a linear expectation anchored to the bucket's share of the minimal floor and fitted on the low-N half — 1.0 = grew exactly linearly, >1 bends up. The super-linearity lives where × lin (and k) are red.

buckett@N=25t@N=200Δ msshare× lin∝Nk
generateOutput (self)26307+28031%1.48×1.21
specializeModule37295+25729%1.06×0.99
simplifyIR13133+12013%1.34×1.11
linkAndOptimizeIR (self)1193+839%1.19×1.07
SemanticChecking2988+597%0.60×0.60

Also growing (below top-5): generateIR (+33 ms, 4%), legalizeExistentialTypeLayout (+26 ms, 3%), legalizeResourceTypes (+25 ms, 3%).

Near-constant (≤2% of growth each): linkIR (3→6 ms), performMandatoryEarlyInlining (1→5 ms), performForceInlining (0→3 ms), compileInner (self) (0→1 ms), unrollLoopsInModule (0→1 ms), parseTranslationUnit (0→1 ms), frontEndExecute (self) (0→1 ms).

Sweep numbers (median ms)

NcompileInnerlinkAndOptimizeIRfrontEndExecute
251366841
5023412555
10045725681
2001032589135