Mindverse Research / In development

One call.
The right model.

Hydra matches your task to the right model. Quality, cost and response time guide selection — with a traceable decision for every request.

Currently: local development planner. Production access is not yet activated.

Illustrative model selection

Your request
hydra

Task + priorities + eligible models

Model A
Model B ✓
Model C

One answer. One recorded choice.

Concept illustration. No request is sent and no model is executed.

Model pool
25+
Models in the current development pool
Internal benchmark
85.43 / 100
Hydra adjustment D · September 2026
Generation cost
$0.02062
Per 100 planned tasks · excludes service overhead

// 01 · Selection

Different tasks. Different models.

One model per request. Your task and priorities determine which one.

Model A

Extract invoice data

Return invoice number, vendor and total as JSON.

Structured output and the quality requirement narrow the selection.

Model B

Review code

Check a retry condition for a boundary error.

One model with suitable reasoning capability analyses the code and returns its finding.

Model C

Draft a customer reply

Respond empathetically to a delayed delivery without unsupported promises.

A writing-capable model works with the supplied facts and tone requirements.

Illustrative examples with placeholder models. No live requests; the historical benchmark below is separate.

// 02 · How selection works

A route with a reason.

01 / Understand

The task sets the direction.

Task type, reasoning needs and the requested output format shape which models suit a request.

02 / Select

Your priorities guide the choice.

The current planner combines application-supplied model eligibility with development measurements for the task category. Quality, generation cost and response time can be weighted.

03 / Use one model

One request. One selected model.

The illustrated route uses one model for a direct response. Other models remain alternatives. Multi-model chains are not currently available.

04 / Trace

A decision with a reason.

Task features, selection criteria and chosen models can be traced in the local decision record.

Coordinator in Germany

Routing takes place in Germany. Selected models may run with external providers.

Open-model filter

An optional portfolio filter narrows selection using reviewed model classifications and explicit self-hosting assumptions. Model weights and licence terms need to be checked before deployment.

Privacy as a requirement

Your privacy requirements shape which routes are eligible. The current planner uses the eligibility supplied by the application.

25+ models in one pool

The development pool currently covers 29 models and 261 task combinations, with task-specific rankings and adjustable priorities.

// 03 · Internal benchmark

The numbers. With their limits.

Three Hydra configurations and six historical reference models on 100 internal tasks, evaluated with Fable 5.1. Development results from September 2026.

9 selected configurations · sorted by mean quality
ConfigurationScored / 100Quality / 100USD / 100 planned tasks
GLM 5.3 FlashHistorical reference model99
87.95
$0.11406
DeepSeek V4 Pro 0813Historical reference model99
87.72
$0.39692
Hydra adjustment DHydra configuration98
85.43
$0.02062
GPT-5.6 Luna maxHistorical reference model99
84.43
$0.18695
Hydra adjustment AHydra configuration99
84.06
≥ $0.03989 + ?
MiniMax M3Historical reference model99
83.91
$0.20855
Hydra adjustment H · OSSHydra configuration95
83.83
≥ $0.04837 + ?
GPT-OSS 120B highHistorical reference model95
82.82
≥ $0.14708 + ?
GPT-5.6 Luna lowHistorical reference model99
81.77
$0.02974

Quality

Mean Fable 5.1 score over scored tasks, not all 100 planned tasks. The source chart shows descriptive 95% task-bootstrap intervals (2,000 replicates). The task set was reused; there is no independent holdout evidence.

Run conditions

A and fixed controls: 9 September; D and H: 10 September 2026. Providers, timeouts, recovery and judging differed between waves.

Costs

Reported generation costs in USD for 100 planned tasks, including recorded attempts. Excludes the coordinator, orchestration, judging and service overhead. ≥ + ? marks a lower bound with an unknown remainder, not a complete price.

Comparison limits

Nine selected configurations on a direct-response workload without tools. Unequal scored coverage, historical controls and overlapping intervals do not establish universal savings, superiority or a Pareto frontier. These are internal scores, not Artificial Analysis rankings.

04 · Evaluation & integration

Your workload. Your benchmark.

The current development version plans model selection locally. Let’s discuss the quality, cost and response times your application needs — and what a suitable evaluation should demonstrate.

Node.js / Offline

node current/head.mjs inspect
node current/head.mjs verify

Run from a local Hydra installation. These commands inspect the model pool and verify the planner offline; they do not execute models. Production access is not yet activated.

// FAQ

Frequently asked questions about Hydra

What is Hydra today?
Hydra is a routing project by Mindverse, currently in development. The current planner locally selects a model using task features, eligible models and weighted priorities. Production access is not yet activated.
Does Hydra use several models per request?
Each illustrated request currently uses one selected model. Multi-model routes with planners, parallel workers or separate verifier models are not currently available.
Can I weight quality, cost and speed?
Yes. The development planner combines your quality, generation-cost and response-time preferences with task-specific development measurements. An optional portfolio filter narrows the pool to open models under the documented classifications and self-hosting assumptions.
Does all data stay in Germany?
Routing takes place in Germany. Selected models may run with external providers. Data residency and privacy requirements therefore need to be considered when making models eligible.
How can I evaluate Hydra?
The current local installation provides offline commands to inspect the model pool and verify routing decisions. These do not execute models. Talk to us about your workload and a suitable evaluation.