Skip to content

Performance over time ​

How ComposableTuringIDModels's benchmark suites have moved across recent revisions: an overall summary across the package first, then one section per suite.

Summary ​

Each benchmark suite's headline timing across recent revisions.

SuiteMedian ratioTrendStatus
AD gradients0.98→ok
Model evaluation1.01→ok
Sampling1.03↗ok
time_to_load1.0→ok

Ratio: latest vs oldest shown revision (1.00 = no change, higher = slower/larger). ⚠ reg = at/above the regression threshold.

Tables below show the most recent 4 revisions, columns labelled by commit date.

AD gradients ​

Time ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
AR latent logjoint/ForwardDiff11.8 ± 21 μs14.1 ± 21 μs16.1 ± 21 μs12.9 ± 21 μs
AR latent logjoint/Mooncake reverse8.07 ± 2.2 μs7.91 ± 2 μs7.91 ± 2.2 μs7.93 ± 2 μs
AR latent logjoint/ReverseDiff (tape)0.0466 ± 0.0019 ms0.0505 ± 0.009 ms0.0509 ± 0.0082 ms0.0466 ± 0.0016 ms
DirectInfections+Poisson posterior/Enzyme reverse5.9 ± 0.49 μs0.0708 ± 0.00078 ms0.0764 ± 0.00094 ms5.92 ± 0.5 μs
DirectInfections+Poisson posterior/ForwardDiff19 ± 24 μs0.214 ± 0.026 ms0.217 ± 0.027 ms18.6 ± 23 μs
DirectInfections+Poisson posterior/Mooncake reverse9.44 ± 0.44 μs0.0786 ± 0.00097 ms0.0788 ± 0.001 ms9.45 ± 0.5 μs
DirectInfections+Poisson posterior/ReverseDiff (tape)0.122 ± 0.023 ms0.175 ± 0.024 ms0.175 ± 0.024 ms0.123 ± 0.023 ms

Memory ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
AR latent logjoint/ForwardDiff0.059 k allocs: 0.0512 MB0.056 k allocs: 0.0508 MB0.056 k allocs: 0.0508 MB0.059 k allocs: 0.0512 MB
AR latent logjoint/Mooncake reverse0.043 k allocs: 4.64 kB0.036 k allocs: 4.42 kB0.036 k allocs: 4.42 kB0.043 k allocs: 4.64 kB
AR latent logjoint/ReverseDiff (tape)0.738 k allocs: 30.6 kB0.775 k allocs: 0.0319 MB0.775 k allocs: 0.0319 MB0.738 k allocs: 30.6 kB
DirectInfections+Poisson posterior/Enzyme reverse0.034 k allocs: 4.36 kB0.242 k allocs: 12.3 kB0.242 k allocs: 12.3 kB0.035 k allocs: 4.39 kB
DirectInfections+Poisson posterior/ForwardDiff0.068 k allocs: 0.0591 MB0.68 k allocs: 0.0835 MB0.68 k allocs: 0.0835 MB0.068 k allocs: 0.0591 MB
DirectInfections+Poisson posterior/Mooncake reverse0.038 k allocs: 4.83 kB0.316 k allocs: 14.9 kB0.316 k allocs: 14.9 kB0.038 k allocs: 4.83 kB
DirectInfections+Poisson posterior/ReverseDiff (tape)1.7 k allocs: 0.0643 MB1.65 k allocs: 0.0654 MB1.65 k allocs: 0.0654 MB1.7 k allocs: 0.0643 MB

Model evaluation ​

Time ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
AR latent/forward0.665 ± 1 μs0.672 ± 0.13 μs0.654 ± 0.16 μs0.657 ± 1 μs
AR latent/rand1.94 ± 1.2 μs1.87 ± 1.2 μs0.873 ± 1.2 μs1.95 ± 1.2 μs
DirectInfections+Poisson/forward2.27 ± 0.97 μs0.0676 ± 0.00089 ms0.068 ± 0.00085 ms2.29 ± 0.96 μs
DirectInfections+Poisson/rand1.98 ± 1.1 μs0.0666 ± 0.00082 ms0.0666 ± 0.00082 ms2 ± 1.2 μs
RandomWalk latent/forward1.28 ± 0.7 μs1.27 ± 0.71 μs1.25 ± 0.7 μs1.3 ± 0.64 μs
RandomWalk latent/rand1.44 ± 0.93 μs1.46 ± 0.95 μs1.45 ± 0.93 μs1.46 ± 0.95 μs
Renewal+NegativeBinomial/forward7.68 ± 2.9 μs0.0728 ± 0.00097 ms0.0726 ± 0.00094 ms7.59 ± 2.8 μs
Renewal+NegativeBinomial/rand5.21 ± 3 μs0.0698 ± 0.0012 ms0.0702 ± 0.0012 ms5.19 ± 3.1 μs

Memory ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
AR latent/forward21 allocs: 2.44 kB20 allocs: 2.41 kB20 allocs: 2.41 kB21 allocs: 2.44 kB
AR latent/rand23 allocs: 2.86 kB22 allocs: 2.83 kB22 allocs: 2.83 kB23 allocs: 2.86 kB
DirectInfections+Poisson/forward23 allocs: 2.52 kB0.35 k allocs: 15.8 kB0.35 k allocs: 15.8 kB23 allocs: 2.52 kB
DirectInfections+Poisson/rand20 allocs: 2.67 kB0.349 k allocs: 15.1 kB0.349 k allocs: 15.1 kB20 allocs: 2.67 kB
RandomWalk latent/forward17 allocs: 1.86 kB16 allocs: 1.83 kB16 allocs: 1.83 kB17 allocs: 1.86 kB
RandomWalk latent/rand16 allocs: 2.08 kB15 allocs: 2.05 kB15 allocs: 2.05 kB16 allocs: 2.08 kB
Renewal+NegativeBinomial/forward0.153 k allocs: 8.28 kB0.57 k allocs: 23.7 kB0.57 k allocs: 23.7 kB0.153 k allocs: 8.28 kB
Renewal+NegativeBinomial/rand0.148 k allocs: 8.38 kB0.567 k allocs: 23 kB0.567 k allocs: 23 kB0.148 k allocs: 8.38 kB

Sampling ​

Time ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
NUTS (DirectInfections+Poisson, 50 draws)0.128 ± 0.014 s0.962 ± 0.041 s0.973 ± 0.033 s0.132 ± 0.0056 s

Memory ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
NUTS (DirectInfections+Poisson, 50 draws)0.479 M allocs: 0.269 GB2.99 M allocs: 0.371 GB2.99 M allocs: 0.371 GB0.479 M allocs: 0.269 GB

time_to_load ​

Time ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
time_to_load4.91 ± 0.11 s4.91 ± 0.051 s5.12 ± 0.014 s4.91 ± 0.17 s

Memory ​

Benchmarkv0.1.2v0.1.1v0.1.02026-09-07
time_to_load0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB0.149 k allocs: 11.2 kB

Per-benchmark timelines ​

<details> <summary>Show 1 plot</summary>

plot_ComposableTuringIDModels.png

</details>

About these benchmarks ​

ComposableTuringIDModels tracks the performance of representative modelling operations over time. The suite is a prototype: it covers a small set of representative models rather than exhaustively measuring every component, and the numbers are indicative rather than a guarantee.

Benchmarking reuses the shared tooling in EpiAwarePackageTools.Benchmarks rather than re-implementing a runner or a comparison report. The package owns the suite definition (benchmark/benchmarks.jl); the kit owns running it and turning results into a legible pull-request comment.

What is measured ​

The suite is a BenchmarkTools.BenchmarkGroup named SUITE, defined in benchmark/benchmarks.jl, with three groups.

  • Model evaluation — building and evaluating representative models. For each model the suite times a prior draw (rand, which samples every random variable) and the forward pass (model(), which returns the generated quantities). The models are two latent processes (AR, RandomWalk) and two composed IDModels (DirectInfections with PoissonError, and Renewal with NegativeBinomialError), each turned into a Turing model via as_turing_model.

  • Sampling — a short NUTS run (50 draws) on a composed DirectInfections + PoissonError model conditioned on data simulated from its own prior.

  • AD gradients — the gradient of a representative log-density across automatic-differentiation backends (ForwardDiff, ReverseDiff, Mooncake, and Enzyme where supported). Results are keyed by scenario and backend so the comparison report folds them into a per-(scenario × backend) matrix.

Running the suite locally ​

The benchmark/ directory is its own Julia environment. Run the whole suite and save the results with the managed runner, which calls EpiAwarePackageTools.Benchmarks.run_suite:

sh
julia --project=benchmark benchmark/run.jl results.json

run_suite uses a short per-benchmark time budget so a full run stays affordable while the minimum-time estimator used in the comparison stays stable. To compare two result files and write a Markdown report, use the managed comparison script, which calls EpiAwarePackageTools.Benchmarks.compare_comment:

sh
julia --project=benchmark benchmark/compare.jl pr.json base.json comment.md

Continuous integration ​

Two workflows drive benchmarking in CI, both building on the shared kit.

  • benchmark.yaml runs on pull requests. It benchmarks the pull-request head and the base branch in separate jobs, then posts (and updates) a single comparison comment: a bucketed summary plus collapsed per-benchmark tables split into evaluation and AD-gradient groups.

  • benchmark-history.yaml runs on pushes to main and on tags. It benchmarks the recent tagged releases plus the current commit with AirspeedVelocity and publishes a timeline to the repository's benchmarks branch. The kit's asv_comment / flatten_asv helpers read the same AirspeedVelocity result format when a report is needed.