Performance over time
How ComposableTuringIDModels's benchmark suites have moved across recent revisions: an overall summary across the package first, then one section per suite.
Summary
Each benchmark suite's headline timing across recent revisions.
| Suite | Median ratio | Trend | Status |
|---|---|---|---|
| AD gradients | 1.0 | → | ok |
| Model evaluation | 0.71 | ↘ | ok |
| Sampling | 0.99 | → | ok |
| time_to_load | 1.01 | → | ok |
Ratio: latest vs oldest shown revision (1.00 = no change, higher = slower/larger). ⚠ reg = at/above the regression threshold.

Tables below show the most recent 4 revisions, columns labelled by commit date.
AD gradients
Time
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| AR latent logjoint/ForwardDiff | 8.91 ± 10 μs | 8.78 ± 9.5 μs | 9.14 ± 9.9 μs | 8.87 ± 10 μs |
| AR latent logjoint/Mooncake reverse | 4.9 ± 0.42 μs | 4.85 ± 0.43 μs | 4.9 ± 0.39 μs | 4.96 ± 0.41 μs |
| AR latent logjoint/ReverseDiff (tape) | 24 ± 5 μs | 26.4 ± 5.7 μs | 26.4 ± 5.4 μs | 24.2 ± 1.5 μs |
| DirectInfections+Poisson posterior/Enzyme reverse | 2.68 ± 0.28 μs | 0.0378 ± 0.00053 ms | 0.0376 ± 0.00047 ms | 2.7 ± 0.28 μs |
| DirectInfections+Poisson posterior/ForwardDiff | 13.1 ± 12 μs | 0.116 ± 0.014 ms | 0.114 ± 0.014 ms | 13.4 ± 12 μs |
| DirectInfections+Poisson posterior/Mooncake reverse | 6.47 ± 0.46 μs | 0.0422 ± 0.00081 ms | 0.0422 ± 0.00073 ms | 6.48 ± 0.45 μs |
| DirectInfections+Poisson posterior/ReverseDiff (tape) | 0.066 ± 0.013 ms | 0.0935 ± 0.015 ms | 0.0935 ± 0.015 ms | 0.0662 ± 0.015 ms |
Memory
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| AR latent logjoint/ForwardDiff | 0.059 k allocs: 0.0512 MB | 0.056 k allocs: 0.0508 MB | 0.056 k allocs: 0.0508 MB | 0.059 k allocs: 0.0512 MB |
| AR latent logjoint/Mooncake reverse | 0.047 k allocs: 5.2 kB | 0.04 k allocs: 4.98 kB | 0.04 k allocs: 4.98 kB | 0.047 k allocs: 5.2 kB |
| AR latent logjoint/ReverseDiff (tape) | 0.738 k allocs: 30.6 kB | 0.775 k allocs: 0.0319 MB | 0.775 k allocs: 0.0319 MB | 0.738 k allocs: 30.6 kB |
| DirectInfections+Poisson posterior/Enzyme reverse | 0.034 k allocs: 4.36 kB | 0.242 k allocs: 12.3 kB | 0.242 k allocs: 12.3 kB | 0.034 k allocs: 4.36 kB |
| DirectInfections+Poisson posterior/ForwardDiff | 0.068 k allocs: 0.0591 MB | 0.68 k allocs: 0.0835 MB | 0.68 k allocs: 0.0835 MB | 0.068 k allocs: 0.0591 MB |
| DirectInfections+Poisson posterior/Mooncake reverse | 0.042 k allocs: 5.39 kB | 0.32 k allocs: 15.5 kB | 0.32 k allocs: 15.5 kB | 0.042 k allocs: 5.39 kB |
| DirectInfections+Poisson posterior/ReverseDiff (tape) | 1.7 k allocs: 0.0643 MB | 1.65 k allocs: 0.0654 MB | 1.65 k allocs: 0.0654 MB | 1.7 k allocs: 0.0643 MB |
Model evaluation
Time
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| AR latent/forward | 0.383 ± 0.57 μs | 0.363 ± 0.035 μs | 0.367 ± 0.036 μs | 0.395 ± 0.59 μs |
| AR latent/rand | 1.07 ± 0.73 μs | 0.474 ± 0.71 μs | 0.513 ± 0.72 μs | 1.02 ± 0.73 μs |
| DirectInfections+Poisson/forward | 1.31 ± 0.62 μs | 0.0358 ± 0.00051 ms | 0.036 ± 0.00074 ms | 1.32 ± 0.61 μs |
| DirectInfections+Poisson/rand | 1.07 ± 0.72 μs | 0.0352 ± 0.00058 ms | 0.0352 ± 0.00053 ms | 0.504 ± 0.72 μs |
| RandomWalk latent/forward | 0.315 ± 0.44 μs | 0.301 ± 0.44 μs | 0.302 ± 0.44 μs | 0.328 ± 0.44 μs |
| RandomWalk latent/rand | 0.331 ± 0.54 μs | 0.302 ± 0.54 μs | 0.317 ± 0.54 μs | 0.351 ± 0.54 μs |
| Renewal+NegativeBinomial/forward | 4.19 ± 1.6 μs | 0.0386 ± 0.00068 ms | 0.0383 ± 0.00065 ms | 4.2 ± 1.6 μs |
| Renewal+NegativeBinomial/rand | 2.65 ± 1.7 μs | 0.0371 ± 0.00072 ms | 0.0369 ± 0.00082 ms | 2.6 ± 1.6 μs |
Memory
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| AR latent/forward | 21 allocs: 2.44 kB | 20 allocs: 2.41 kB | 20 allocs: 2.41 kB | 21 allocs: 2.44 kB |
| AR latent/rand | 23 allocs: 2.86 kB | 22 allocs: 2.83 kB | 22 allocs: 2.83 kB | 23 allocs: 2.86 kB |
| DirectInfections+Poisson/forward | 23 allocs: 2.52 kB | 0.35 k allocs: 15.8 kB | 0.35 k allocs: 15.8 kB | 23 allocs: 2.52 kB |
| DirectInfections+Poisson/rand | 20 allocs: 2.67 kB | 0.349 k allocs: 15.1 kB | 0.349 k allocs: 15.1 kB | 20 allocs: 2.67 kB |
| RandomWalk latent/forward | 17 allocs: 1.86 kB | 16 allocs: 1.83 kB | 16 allocs: 1.83 kB | 17 allocs: 1.86 kB |
| RandomWalk latent/rand | 16 allocs: 2.08 kB | 15 allocs: 2.05 kB | 15 allocs: 2.05 kB | 16 allocs: 2.08 kB |
| Renewal+NegativeBinomial/forward | 0.153 k allocs: 8.28 kB | 0.57 k allocs: 23.7 kB | 0.57 k allocs: 23.7 kB | 0.153 k allocs: 8.28 kB |
| Renewal+NegativeBinomial/rand | 0.148 k allocs: 8.38 kB | 0.567 k allocs: 23 kB | 0.567 k allocs: 23 kB | 0.148 k allocs: 8.38 kB |
Sampling
Time
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| NUTS (DirectInfections+Poisson, 50 draws) | 0.0771 ± 0.0032 s | 0.528 ± 0.0083 s | 0.527 ± 0.018 s | 0.0765 ± 0.0037 s |
Memory
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| NUTS (DirectInfections+Poisson, 50 draws) | 0.479 M allocs: 0.269 GB | 2.99 M allocs: 0.371 GB | 2.99 M allocs: 0.371 GB | 0.479 M allocs: 0.269 GB |
time_to_load
Time
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| time_to_load | 3.37 ± 0.024 s | 2.98 ± 0.018 s | 2.98 ± 0.012 s | 3.4 ± 0.023 s |
Memory
| Benchmark | v0.1.2 | v0.1.1 | v0.1.0 | 1a3e4853454310... |
|---|---|---|---|---|
| time_to_load | 0.149 k allocs: 11.2 kB | 0.149 k allocs: 11.2 kB | 0.149 k allocs: 11.2 kB | 0.149 k allocs: 11.2 kB |
Per-benchmark timelines
<details> <summary>Show 1 plot</summary>

</details>
About these benchmarks
ComposableTuringIDModels tracks the performance of representative modelling operations over time. The suite is a prototype: it covers a small set of representative models rather than exhaustively measuring every component, and the numbers are indicative rather than a guarantee.
Benchmarking reuses the shared tooling in EpiAwarePackageTools.Benchmarks rather than re-implementing a runner or a comparison report. The package owns the suite definition (benchmark/benchmarks.jl); the kit owns running it and turning results into a legible pull-request comment.
What is measured
The suite is a BenchmarkTools.BenchmarkGroup named SUITE, defined in benchmark/benchmarks.jl, with three groups.
Model evaluation — building and evaluating representative models. For each model the suite times a prior draw (
rand, which samples every random variable) and the forward pass (model(), which returns the generated quantities). The models are two latent processes (AR,RandomWalk) and two composedIDModels (DirectInfectionswithPoissonError, andRenewalwithNegativeBinomialError), each turned into a Turing model viaas_turing_model.Sampling — a short NUTS run (50 draws) on a composed
DirectInfections+PoissonErrormodel conditioned on data simulated from its own prior.AD gradients — the gradient of a representative log-density across automatic-differentiation backends (
ForwardDiff,ReverseDiff,Mooncake, andEnzymewhere supported). Results are keyed by scenario and backend so the comparison report folds them into a per-(scenario × backend) matrix.
Running the suite locally
The benchmark/ directory is its own Julia environment. Run the whole suite and save the results with the managed runner, which calls EpiAwarePackageTools.Benchmarks.run_suite:
julia --project=benchmark benchmark/run.jl results.jsonrun_suite uses a short per-benchmark time budget so a full run stays affordable while the minimum-time estimator used in the comparison stays stable. To compare two result files and write a Markdown report, use the managed comparison script, which calls EpiAwarePackageTools.Benchmarks.compare_comment:
julia --project=benchmark benchmark/compare.jl pr.json base.json comment.mdContinuous integration
Two workflows drive benchmarking in CI, both building on the shared kit.
benchmark.yamlruns on pull requests. It benchmarks the pull-request head and the base branch in separate jobs, then posts (and updates) a single comparison comment: a bucketed summary plus collapsed per-benchmark tables split into evaluation and AD-gradient groups.benchmark-history.yamlruns on pushes tomainand on tags. It benchmarks the recent tagged releases plus the current commit with AirspeedVelocity and publishes a timeline to the repository'sbenchmarksbranch. The kit'sasv_comment/flatten_asvhelpers read the same AirspeedVelocity result format when a report is needed.