Skip to content

Performance over time

How ComposableTuringIDModels's benchmark suites have moved across recent revisions: an overall summary across the package first, then one section per suite.

Summary

Each benchmark suite's headline timing across recent revisions.

SuiteMedian ratioTrendStatus
AD gradients1.0ok
Model evaluation0.71ok
Sampling0.99ok
time_to_load1.01ok

Ratio: latest vs oldest shown revision (1.00 = no change, higher = slower/larger). ⚠ reg = at/above the regression threshold.

Tables below show the most recent 4 revisions, columns labelled by commit date.

AD gradients

Time

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
AR latent logjoint/ForwardDiff8.91 ± 10 μs8.78 ± 9.5 μs9.14 ± 9.9 μs8.87 ± 10 μs
AR latent logjoint/Mooncake reverse4.9 ± 0.42 μs4.85 ± 0.43 μs4.9 ± 0.39 μs4.96 ± 0.41 μs
AR latent logjoint/ReverseDiff (tape)24 ± 5 μs26.4 ± 5.7 μs26.4 ± 5.4 μs24.2 ± 1.5 μs
DirectInfections+Poisson posterior/Enzyme reverse2.68 ± 0.28 μs0.0378 ± 0.00053 ms0.0376 ± 0.00047 ms2.7 ± 0.28 μs
DirectInfections+Poisson posterior/ForwardDiff13.1 ± 12 μs0.116 ± 0.014 ms0.114 ± 0.014 ms13.4 ± 12 μs
DirectInfections+Poisson posterior/Mooncake reverse6.47 ± 0.46 μs0.0422 ± 0.00081 ms0.0422 ± 0.00073 ms6.48 ± 0.45 μs
DirectInfections+Poisson posterior/ReverseDiff (tape)0.066 ± 0.013 ms0.0935 ± 0.015 ms0.0935 ± 0.015 ms0.0662 ± 0.015 ms

Memory

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
AR latent logjoint/ForwardDiff0.059 k allocs: 0.0512 MB0.056 k allocs: 0.0508 MB0.056 k allocs: 0.0508 MB0.059 k allocs: 0.0512 MB
AR latent logjoint/Mooncake reverse0.047 k allocs: 5.2 kB0.04 k allocs: 4.98 kB0.04 k allocs: 4.98 kB0.047 k allocs: 5.2 kB
AR latent logjoint/ReverseDiff (tape)0.738 k allocs: 30.6 kB0.775 k allocs: 0.0319 MB0.775 k allocs: 0.0319 MB0.738 k allocs: 30.6 kB
DirectInfections+Poisson posterior/Enzyme reverse0.034 k allocs: 4.36 kB0.242 k allocs: 12.3 kB0.242 k allocs: 12.3 kB0.034 k allocs: 4.36 kB
DirectInfections+Poisson posterior/ForwardDiff0.068 k allocs: 0.0591 MB0.68 k allocs: 0.0835 MB0.68 k allocs: 0.0835 MB0.068 k allocs: 0.0591 MB
DirectInfections+Poisson posterior/Mooncake reverse0.042 k allocs: 5.39 kB0.32 k allocs: 15.5 kB0.32 k allocs: 15.5 kB0.042 k allocs: 5.39 kB
DirectInfections+Poisson posterior/ReverseDiff (tape)1.7 k allocs: 0.0643 MB1.65 k allocs: 0.0654 MB1.65 k allocs: 0.0654 MB1.7 k allocs: 0.0643 MB

Model evaluation

Time

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
AR latent/forward0.383 ± 0.57 μs0.363 ± 0.035 μs0.367 ± 0.036 μs0.395 ± 0.59 μs
AR latent/rand1.07 ± 0.73 μs0.474 ± 0.71 μs0.513 ± 0.72 μs1.02 ± 0.73 μs
DirectInfections+Poisson/forward1.31 ± 0.62 μs0.0358 ± 0.00051 ms0.036 ± 0.00074 ms1.32 ± 0.61 μs
DirectInfections+Poisson/rand1.07 ± 0.72 μs0.0352 ± 0.00058 ms0.0352 ± 0.00053 ms0.504 ± 0.72 μs
RandomWalk latent/forward0.315 ± 0.44 μs0.301 ± 0.44 μs0.302 ± 0.44 μs0.328 ± 0.44 μs
RandomWalk latent/rand0.331 ± 0.54 μs0.302 ± 0.54 μs0.317 ± 0.54 μs0.351 ± 0.54 μs
Renewal+NegativeBinomial/forward4.19 ± 1.6 μs0.0386 ± 0.00068 ms0.0383 ± 0.00065 ms4.2 ± 1.6 μs
Renewal+NegativeBinomial/rand2.65 ± 1.7 μs0.0371 ± 0.00072 ms0.0369 ± 0.00082 ms2.6 ± 1.6 μs

Memory

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
AR latent/forward21 allocs: 2.44 kB20 allocs: 2.41 kB20 allocs: 2.41 kB21 allocs: 2.44 kB
AR latent/rand23 allocs: 2.86 kB22 allocs: 2.83 kB22 allocs: 2.83 kB23 allocs: 2.86 kB
DirectInfections+Poisson/forward23 allocs: 2.52 kB0.35 k allocs: 15.8 kB0.35 k allocs: 15.8 kB23 allocs: 2.52 kB
DirectInfections+Poisson/rand20 allocs: 2.67 kB0.349 k allocs: 15.1 kB0.349 k allocs: 15.1 kB20 allocs: 2.67 kB
RandomWalk latent/forward17 allocs: 1.86 kB16 allocs: 1.83 kB16 allocs: 1.83 kB17 allocs: 1.86 kB
RandomWalk latent/rand16 allocs: 2.08 kB15 allocs: 2.05 kB15 allocs: 2.05 kB16 allocs: 2.08 kB
Renewal+NegativeBinomial/forward0.153 k allocs: 8.28 kB0.57 k allocs: 23.7 kB0.57 k allocs: 23.7 kB0.153 k allocs: 8.28 kB
Renewal+NegativeBinomial/rand0.148 k allocs: 8.38 kB0.567 k allocs: 23 kB0.567 k allocs: 23 kB0.148 k allocs: 8.38 kB

Sampling

Time

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
NUTS (DirectInfections+Poisson, 50 draws)0.0771 ± 0.0032 s0.528 ± 0.0083 s0.527 ± 0.018 s0.0765 ± 0.0037 s

Memory

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
NUTS (DirectInfections+Poisson, 50 draws)0.479 M allocs: 0.269 GB2.99 M allocs: 0.371 GB2.99 M allocs: 0.371 GB0.479 M allocs: 0.269 GB

time_to_load

Time

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
time_to_load3.37 ± 0.024 s2.98 ± 0.018 s2.98 ± 0.012 s3.4 ± 0.023 s

Memory

Benchmarkv0.1.2v0.1.1v0.1.01a3e4853454310...
time_to_load0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB

Per-benchmark timelines

<details> <summary>Show 1 plot</summary>

plot_ComposableTuringIDModels.png

</details>

About these benchmarks

ComposableTuringIDModels tracks the performance of representative modelling operations over time. The suite is a prototype: it covers a small set of representative models rather than exhaustively measuring every component, and the numbers are indicative rather than a guarantee.

Benchmarking reuses the shared tooling in EpiAwarePackageTools.Benchmarks rather than re-implementing a runner or a comparison report. The package owns the suite definition (benchmark/benchmarks.jl); the kit owns running it and turning results into a legible pull-request comment.

What is measured

The suite is a BenchmarkTools.BenchmarkGroup named SUITE, defined in benchmark/benchmarks.jl, with three groups.

  • Model evaluation — building and evaluating representative models. For each model the suite times a prior draw (rand, which samples every random variable) and the forward pass (model(), which returns the generated quantities). The models are two latent processes (AR, RandomWalk) and two composed IDModels (DirectInfections with PoissonError, and Renewal with NegativeBinomialError), each turned into a Turing model via as_turing_model.

  • Sampling — a short NUTS run (50 draws) on a composed DirectInfections + PoissonError model conditioned on data simulated from its own prior.

  • AD gradients — the gradient of a representative log-density across automatic-differentiation backends (ForwardDiff, ReverseDiff, Mooncake, and Enzyme where supported). Results are keyed by scenario and backend so the comparison report folds them into a per-(scenario × backend) matrix.

Running the suite locally

The benchmark/ directory is its own Julia environment. Run the whole suite and save the results with the managed runner, which calls EpiAwarePackageTools.Benchmarks.run_suite:

sh
julia --project=benchmark benchmark/run.jl results.json

run_suite uses a short per-benchmark time budget so a full run stays affordable while the minimum-time estimator used in the comparison stays stable. To compare two result files and write a Markdown report, use the managed comparison script, which calls EpiAwarePackageTools.Benchmarks.compare_comment:

sh
julia --project=benchmark benchmark/compare.jl pr.json base.json comment.md

Continuous integration

Two workflows drive benchmarking in CI, both building on the shared kit.

  • benchmark.yaml runs on pull requests. It benchmarks the pull-request head and the base branch in separate jobs, then posts (and updates) a single comparison comment: a bucketed summary plus collapsed per-benchmark tables split into evaluation and AD-gradient groups.

  • benchmark-history.yaml runs on pushes to main and on tags. It benchmarks the recent tagged releases plus the current commit with AirspeedVelocity and publishes a timeline to the repository's benchmarks branch. The kit's asv_comment / flatten_asv helpers read the same AirspeedVelocity result format when a report is needed.