Data scaling
length 60, 8 states, 12 iterations| sequences | observations | hmmlearn | mixle | speedup |
|---|---|---|---|---|
| 250 | 15,000 | 352 ms | 23 ms | 15.5x |
| 1,000 | 60,000 | 1,405 ms | 79 ms | 17.8x |
| 4,000 | 240,000 | 5,695 ms | 316 ms | 18.0x |
Benchmarks
Fit-time scaling results for full-covariance Gaussian mixtures, Gaussian-emission HMMs, GPU fits, and profiling runs. Every comparison uses shared data, shared initialization, fixed iteration counts, and final log-likelihood checks.
Methodology
The harness measures comparable work: the same input arrays, the same initial parameters, and the same number of EM or Baum-Welch iterations.
Gaussian HMM
Many sequences from a latent regime process, Gaussian emissions, shared initialization, and fixed Baum-Welch iterations.
| sequences | observations | hmmlearn | mixle | speedup |
|---|---|---|---|---|
| 250 | 15,000 | 352 ms | 23 ms | 15.5x |
| 1,000 | 60,000 | 1,405 ms | 79 ms | 17.8x |
| 4,000 | 240,000 | 5,695 ms | 316 ms | 18.0x |
| states | hmmlearn | mixle | speedup |
|---|---|---|---|
| 4 | 876 ms | 45 ms | 19.4x |
| 8 | 1,403 ms | 78 ms | 17.9x |
| 16 | 3,448 ms | 158 ms | 21.8x |
| 32 | 10,886 ms | 435 ms | 25.0x |
Full-covariance GMM
scikit-learn and pomegranate are faster on small fits. mixle carries more fixed overhead, but its covariance accumulation scales better and crosses over near N=140k.
| N | sklearn | pomegranate | mixle | vs sklearn |
|---|---|---|---|---|
| 5,000 | 101 ms | 125 ms | 144 ms | 0.70x |
| 20,000 | 399 ms | 316 ms | 615 ms | 0.65x |
| 80,000 | 2,227 ms | 2,318 ms | 2,870 ms | 0.78x |
| 200,000 | 5,522 ms | 5,812 ms | 3,919 ms | 1.41x |
| dim | sklearn | pomegranate | mixle |
|---|---|---|---|
| 8 | 85 ms | 57 ms | 89 ms |
| 16 | 105 ms | 71 ms | 122 ms |
| 32 | 170 ms | 152 ms | 241 ms |
| 64 | 384 ms | 362 ms | 540 ms |
| 128 | 835 ms | 881 ms | 1,206 ms |
The stable crossover check measured mixle at 0.78x at N=80k, then 1.42x / 1.42x / 1.41x at 140k / 200k / 350k, with the same final likelihood.
GPU engine
This is a capability panel, not a cross-machine speedup claim. The data-scaling run used an RTX 2080 Ti at dim 64 with 16 components.
| N | fit time | peak GPU memory |
|---|---|---|
| 50,000 | 3.3 s | 1.7 GB |
| 200,000 | 8.9 s | 6.8 GB |
| 500,000 | 20.2 s | 16.9 GB |
The dim-128 memory fix reduced a failing 21 GB intermediate to a 2.7 GB peak on an RTX 3060, with CPU parity at 2.5e-13. N=1M at dim 64 still needs chunking over N.
Scaling profile
A separate A4000 + 16-core profiling run measured empirical slopes and cProfile hotspots across model families.