- Home
- Evaluation Engine
Keido Evaluation Engine
Your AI is running. Is it performing?
Your AI is running. Is it performing?
KEE scores every response your AI systems produce, in production, against the four things that decide whether people trust them: speed, accuracy, coverage and reliability. When a score moves, you know the same day, not after the wrong decision has been made.
- Data Readiness
- API
- Visual
- Evaluation Engine
- Care
What it evaluates
If you can log a prompt and a response, KEE can score it.
If you can log a prompt and a response, KEE can score it.
KEE sits alongside any AI system that produces text or structured output: a search or retrieval-augmented assistant, a summariser, a classifier, an agent running multi-step tasks, or a model you have fine-tuned yourself.
Works with
- OpenAI
- Anthropic
- Open-weight models
- Your own models
The four scores
Four scores, each tied to a business consequence.
Four scores, each tied to a business consequence.
-
Speed
Time to first response and time to complete, by task type.
Slow tools get abandoned. Speed is the first adoption signal.
-
Accuracy
Whether the answer is correct, grounded in your sources, and free of fabrication.
The score that decides whether a result can be acted on.
-
Coverage
Whether the answer used everything relevant and left nothing important out.
A correct but incomplete answer still produces a wrong decision.
-
Reliability
How consistently the system performs across conditions, users and edge cases, rolled into one confidence index.
The number a leader can read without knowing the other three.
What you see
One view for leaders. The detail for teams.
One view for leaders. The detail for teams.
-
Leaders
A single view: which systems are performing to expectation and which have moved.
-
Teams
Which queries, which sources, which conditions caused the shift, and a ranked list of what to fix first.
-
Alerts
Configurable by threshold and by audience.
More than a dashboard
-
Reports what happened.
-
Runs the evaluation continuously and tells you what to change.
Which prompts to tighten, which sources to update, which edge cases need a guardrail. As your data grows and demand rises, KEE re-baselines so a score still means the same thing six months in.
Built for governance
Every score is auditable.
Every score is auditable.
- Every score, alert and evaluation run is logged and exportable for audit
- Multi-model and multi-dataset comparison, so you can test a model change before it ships
- Trend analysis by team, task and time, for targeted optimisation
- Runs inside your environment; evaluation data never leaves it
How an engagement runs
Four weeks to a leadership readout.
Four weeks to a leadership readout.
- Week 1 Connect KEE to one system in production and agree the rubrics with the owning team
- Weeks 2 to 3 Baseline scores, tune thresholds, first alert review
- Week 4 Leadership readout with the reliability index and the top five fixes
- Ongoing Monthly or quarterly review, or hand over to your team