Six steps, one Score, five pillars and ten principles.
It's the method we work with and the foundation of BLYNT Platform, our software. We publish it so you know exactly how we reach every number before you trust it.
Every use case goes through six steps
A use case is a specific process you want to improve: sorting complaints, reviewing contracts, preparing reports. They all follow the same path, and none skips a step.
| Step | What happens | What you keep |
|---|---|---|
| 1 · Define | Case summary: problem, goal, owners, monthly volume and human minutes per unit of work. We describe how the system is built and prepare the test suite with its rubric. | Case summary, versioned system and test suite |
| 2 · Baseline | The full suite runs on the current process and is locked as the reference. A person from the client company signs it off. | Signed, immutable baseline |
| 3 · Analyze | The Analyzer applies deterministic rules (no AI, zero tokens) to the baseline, the workflow and the traces, and ranks findings. | Findings with evidence and confidence |
| 4 · Improve | Each useful finding becomes a concrete improvement. The improvement package is tested against the same suite. | Improvements and candidate runs |
| 5 · Validate | The candidate is compared with the baseline against the agreed guardrails: maximum quality drop, critical errors, token and time reduction. If it passes, the company signs off. | Signed validation or a reasoned rejection |
| 6 · Impact | Projection of hours that can be freed and AI spend avoided, with every assumption visible. Case, team and portfolio reports. | PDF report traceable to the source test |
Not using AI yet?
The cycle starts the same way: the manual process is measured and signed off as the baseline. When AI arrives, there is something to compare it with.
Using a licensed chat tool?
If tokens aren't measured, cost leaves the Score and its weight is shared among the other dimensions. This is always flagged.
Who does what?
BLYNT Labs prepares, measures, analyzes and proposes. Your company brings the process knowledge, signs off and decides.
A 0–100 score for how a use case performs
The Score sums up four dimensions with weights set before testing. Changing the weights creates a new policy version and takes earlier runs out of comparison.
| Dimension | Weight | How it's normalized |
|---|---|---|
| Quality | 40% | Average quality according to the suite's rubric |
| Cost | 25% | Cost per accepted result against an agreed ceiling |
| Time | 20% | Processing minutes against a 60-minute ceiling |
| Feedback | 15% | People's average rating, out of 5 |
Score = 0.40·Quality + 0.25·Cost + 0.20·Time + 0.15·Feedback
- Repetitions of each test are averaged first; then each scenario's weight is applied.
- Every dimension is capped between 0 and 100.
- Two runs are only compared if they share the suite, the evaluator and the Score policy.
- The breakdown shows each dimension's weight, value and contribution.
Reproducible diagnosis, with no AI
A deterministic rule engine reads the baseline, the workflow and the per-step traces. It spots, for example, steps using a more expensive model than needed, excessive retries or too many back-and-forth turns to reach the result.
Its signals are hypotheses with evidence, not proven causes. Confidence shows how strong the evidence is, not the odds of success.
Three categories that never mix
- Measured: what came out of the tests.
- Projected: what it would mean at real volume, with the assumptions in view.
- Realized: zero until there is operational evidence.
Hours that can be freed and spend avoided are shown separately. Their sum is "gross opportunity", never ROI or net savings.
Five pillars so the improvement lasts
The Score measures a use case. Maturity measures the organization. They never mix. Each pillar has a checklist of criteria with evidence; its coverage is the percentage met.
AI Engineering & Harness
How the system is built and how it changes without breaking.
Models & Evaluation
Which models are approved and how they are evaluated continuously.
Governance & Security
Documented controls with evidence and history.
People & AI Champions
Who drives adoption and how many people are trained.
Measurement & Continuous Improvement
Whether results are measured, validated and reviewed.
Pillar 1 · AI Engineering & Harness — criteria
- Workflow inventory
- Versioned harness
- Versioned prompts
- Inventoried context sources
- Catalogued MCP connections
- Deterministic steps kept out of the model
- Retry policy
- Context tests
- Rollback procedure
- Change review
Pillar 2 · Models & Evaluation — criteria
- Approved model catalog
- Test suite per use case
- Edge cases covered
- Versioned rubric
- Agreed Score policy
- Baseline per use case
- Documented model comparison
- Evaluation budget
- Periodic evaluation
- Model retirement criteria
Pillar 3 · Governance & Security — criteria
- The organization's governance controls
- Met = documented with evidence
- Every piece of evidence with its history
- The catalog adapts to each company
Pillar 4 · People & AI Champions — criteria
- Designated champions
- Backup champion
- Training path
- Target of trained people
- Quality training
- Community of practice
- Champions in improvements
- Recognition
- Reusable materials
- Adoption tracking
Pillar 5 · Measurement & Continuous Improvement — criteria
- Baseline with a versioned reference
- Indicators with a traceable formula
- Human time measured separately
- AI cost per accepted result
- Projection with visible assumptions
- Measured vs. realized kept apart
- Human validation
- Executive report
- Quarterly review
- Measurement in production
Ten non-negotiable rules
If any part of our work contradicts one of these principles, we fix the work, not the principle.
Measured is not realized
Every number is labeled measured, projected or realized, and categories are never added together.
History is never rewritten
Runs, baselines and validations are immutable. Changing anything creates a new version.
Only compare what is comparable
Same suite, same evaluator, same Score policy. Otherwise, no improvement is calculated.
Success is agreed before testing
Criteria, weights and thresholds are set before anyone sees results.
Explicit human decision
Baseline and validation are signed off by an identified person from the client. Whoever prepares the work doesn't sign it.
Your data stays in your company
BLYNT Platform doesn't call models or connect to your systems. It never stores keys or plain-text prompts.
End-to-end traceability
Every number in the report opens down to its source test. Every action is logged.
Honesty
Whatever can be calculated with code is calculated with code. Informational views are labeled as such.
Client isolation
Each company works in an isolated space. Nobody sees another organization's data.
Specification before code
Nothing is built without an approved specification, and every criterion has its automated test.
The framework, turned into software
Our team works on BLYNT Platform, our own software, which applies the framework step by step. Access is included in every engagement: your company logs in to review, sign off and download reports.
- Versioned use cases, test suites and systems
- Controlled import of results, outputs and traces
- A timer to measure real human time per step
- Analyzer, improvements, validation and PDF reports
- Governance, people, maturity and audit history
- English and Spanish, in each client's currency
Apply the framework to a real process
Start with a four-week assessment and finish with a signed baseline and a written recommendation.