Model A
Extract invoice data
Return invoice number, vendor and total as JSON.
Structured output and the quality requirement narrow the selection.
Mindverse Research / In development
Hydra matches your task to the right model. Quality, cost and response time guide selection — with a traceable decision for every request.
Currently: local development planner. Production access is not yet activated.
Illustrative model selection
Task + priorities + eligible models
One answer. One recorded choice.
Concept illustration. No request is sent and no model is executed.
// 01 · Selection
One model per request. Your task and priorities determine which one.
Model A
Return invoice number, vendor and total as JSON.
Structured output and the quality requirement narrow the selection.
Model B
Check a retry condition for a boundary error.
One model with suitable reasoning capability analyses the code and returns its finding.
Model C
Respond empathetically to a delayed delivery without unsupported promises.
A writing-capable model works with the supplied facts and tone requirements.
Illustrative examples with placeholder models. No live requests; the historical benchmark below is separate.
// 02 · How selection works
01 / Understand
Task type, reasoning needs and the requested output format shape which models suit a request.
02 / Select
The current planner combines application-supplied model eligibility with development measurements for the task category. Quality, generation cost and response time can be weighted.
03 / Use one model
The illustrated route uses one model for a direct response. Other models remain alternatives. Multi-model chains are not currently available.
04 / Trace
Task features, selection criteria and chosen models can be traced in the local decision record.
Routing takes place in Germany. Selected models may run with external providers.
An optional portfolio filter narrows selection using reviewed model classifications and explicit self-hosting assumptions. Model weights and licence terms need to be checked before deployment.
Your privacy requirements shape which routes are eligible. The current planner uses the eligibility supplied by the application.
The development pool currently covers 29 models and 261 task combinations, with task-specific rankings and adjustable priorities.
// 03 · Internal benchmark
Three Hydra configurations and six historical reference models on 100 internal tasks, evaluated with Fable 5.1. Development results from September 2026.
| Configuration | Scored / 100 | Quality / 100 | USD / 100 planned tasks |
|---|---|---|---|
| GLM 5.3 FlashHistorical reference model | 99 | 87.95 | $0.11406 |
| DeepSeek V4 Pro 0813Historical reference model | 99 | 87.72 | $0.39692 |
| Hydra adjustment DHydra configuration | 98 | 85.43 | $0.02062 |
| GPT-5.6 Luna maxHistorical reference model | 99 | 84.43 | $0.18695 |
| Hydra adjustment AHydra configuration | 99 | 84.06 | ≥ $0.03989 + ? |
| MiniMax M3Historical reference model | 99 | 83.91 | $0.20855 |
| Hydra adjustment H · OSSHydra configuration | 95 | 83.83 | ≥ $0.04837 + ? |
| GPT-OSS 120B highHistorical reference model | 95 | 82.82 | ≥ $0.14708 + ? |
| GPT-5.6 Luna lowHistorical reference model | 99 | 81.77 | $0.02974 |
Mean Fable 5.1 score over scored tasks, not all 100 planned tasks. The source chart shows descriptive 95% task-bootstrap intervals (2,000 replicates). The task set was reused; there is no independent holdout evidence.
A and fixed controls: 9 September; D and H: 10 September 2026. Providers, timeouts, recovery and judging differed between waves.
Reported generation costs in USD for 100 planned tasks, including recorded attempts. Excludes the coordinator, orchestration, judging and service overhead. ≥ + ? marks a lower bound with an unknown remainder, not a complete price.
Nine selected configurations on a direct-response workload without tools. Unequal scored coverage, historical controls and overlapping intervals do not establish universal savings, superiority or a Pareto frontier. These are internal scores, not Artificial Analysis rankings.
04 · Evaluation & integration
The current development version plans model selection locally. Let’s discuss the quality, cost and response times your application needs — and what a suitable evaluation should demonstrate.
Node.js / Offline
node current/head.mjs inspect
node current/head.mjs verifyRun from a local Hydra installation. These commands inspect the model pool and verify the planner offline; they do not execute models. Production access is not yet activated.
// FAQ