Decision Workspace
llm-test-bench vs llm-test-bench-core vs serdes-ai-evals
Side-by-side comparison of Rust crates
38
llm-test-bench
experimentalv0.1.0
A production-grade CLI for testing and benchmarking LLM applications with support for GPT-5, Claude Opus 4, Gemini 2.5, and 65+ models
40
llm-test-bench-core
experimentalv0.1.0
Core library for LLM Test Bench - comprehensive testing framework for Large Language Models with 65+ supported models across 14+ providers
49
serdes-ai-evals
experimentalv0.2.6
Evaluation framework for testing and benchmarking serdes-ai agents
Core Metrics
| llm-test-bench | llm-test-bench-core | serdes-ai-evals | |
|---|---|---|---|
| Health Score | 38 | 40 | 49 |
| Total Downloads | 50 | 189 | 735 |
| 30d Downloads | 0 | 0 | 0 |
| Dependents | 0 | 1 | 8 |
| Releases | 1 | 1 | 10 |
| Last Updated | 251d ago | 251d ago | 143d ago |
| Age | 8m | 8m | 5m |
Health Breakdown
llm-test-bench
Maintenance
4
Quality
11
Community
6
Popularity
2
Documentation
15
llm-test-bench-core
Maintenance
4
Quality
11
Community
7
Popularity
3
Documentation
15
serdes-ai-evals
Maintenance
12
Quality
14
Community
8
Popularity
3
Documentation
12
Technical Details
| llm-test-bench | llm-test-bench-core | serdes-ai-evals | |
|---|---|---|---|
| Version | 0.1.0 | 0.1.0 | 0.2.6 |
| Stable (≥1.0) | ✗ No | ✗ No | ✗ No |
| License | MIT OR Apache-2.0 | MIT OR Apache-2.0 | MIT |
| Dependencies | 22 | 55 | 15 |
| Crate Size | 79KB | 355KB | 29KB |
| Features | 0 | 2 | 2 |
| Yanked % | 0.0% | 0.0% | 0.0% |
| Edition | 2021 | 2021 | 2021 |
| MSRV | 1.75.0 | 1.75.0 | 1.75.0 |
| Owners | 1 | 1 | 1 |
Links
Quick Verdict
- •serdes-ai-evals leads with a health score of 49/100, but none of the options score above 80.