Bringing Modern, AI-Driven Quality to Legacy Systems
At-a-Glance
Platforms: LangGraph Multi-Agent System, Web & Mobile Apps, DeepEval, Promptfoo, Langfuse
Core Features: AI routing engine, dynamic agent handoffs, tool-driven automation, metrics tracking & evaluation
Smooth Execution: Synthetic dataset validation, trace-level observability, scalable evaluation pipelines
Delivery Impact: 5,000+ scenarios validated, 2× faster feature delivery driven by stable, measurable AI quality gates
Differentiator: Full-stack AI evaluation system combining DeepEval + Promptfoo + Langfuse with metrics-driven validation

Background
Phoenix AI is building a multi-agent AI platform designed to handle complex, real-world user interactions through dynamic agent orchestration.
Instead of a single chatbot, the system relies on LangGraph-based routing, where a primary agent continuously decides which specialist should respond.
This unlocks powerful capabilities — but also introduces a critical risk:
Without evaluation, the system can silently fail — routing users incorrectly, breaking context, and producing inconsistent outcomes.
We transform this into a controlled, testable, production-ready system.
Goals & Objectives
The objective was to make AI behavior predictable under pressure, measurable, and continuously improvable.
These included:
- Ensure Routing Accuracy: Every input maps to the correct agent
- Control Agent Switching: Eliminate unnecessary or incorrect handoffs
- Validate End-to-End Flows: Ensure reliable execution across workflows and tools
- Test at Scale: Simulate thousands of real-world scenarios
- Establish Quality Gates: Introduce measurable validation before every release
- Track Key Metrics: Monitor routing accuracy, failures, and system behavior over time

Challenges
Multi-agent AI systems fail differently than traditional software — and often invisibly.
1. Invisible Failures
Incorrect routing or outputs don’t crash systems — they degrade trust silently.
2. Non-Deterministic Behavior
The same input produces different outputs, requiring semantic and statistical validation.
3. Routing Is the Core Risk
The system must decide which agent responds — making routing accuracy critical.
4. Fragile Conversation Flow
Uncontrolled switching breaks continuity and reduces usability.
5. Tool & Workflow Reliability
AI doesn’t just respond — it triggers tools and workflows that must execute correctly.
6. No Unified Testing Stack
Validation requires combining multiple tools with clearly defined roles.
Partnership & Process
We built a layered AI evaluation system, where each tool has a clear responsibility and measurable output.
Together, we:
1. DeepEval — Synthetic Dataset & Semantic Evaluation
- Utilise DeepEval to generate and validate 5,000+ synthetic scenarios:
- Realistic user journeys
- Edge cases and ambiguous inputs
- Long conversational flows
- Rare and high-risk situations
- DeepEval enabled semantic scoring and large-scale scenario validation.
2. Promptfoo — Regression & Behavior Testing
- Automated prompt validation
- Defined expected behaviors and scoring rules
- Regression testing across updates
- Output comparison across runs
- Ensuring no regressions reach production.
3. Langfuse — Observability & Metrics Tracking
- Full trace-level visibility
- Agent routing inspection
- Prompt-response debugging
- Metrics tracking (routing accuracy, failure rates, latency, flow success)
- Turning AI into a data-driven system.
4. Tool & Workflow Evaluation
- Tool invocation accuracy
- Cross-system automation reliability
- Input/output consistency
- Failure handling across workflows

Outcomes & Impact
This engagement didn’t just improve quality — it made AI measurable, faster to evolve, and safe to scale.
2× Faster Test Development
Structured evaluation enabled 2× faster delivery of new agentic features (200%+ increase in velocity).
Measurable AI Quality Gates
Every release now passes through quantifiable validation metrics, reducing production risk.
5,000+ Scenarios Validated
DeepEval-powered datasets ensured broad coverage and early detection of failures.
Improved Routing Accuracy
Agent selection became consistent and reliable — minimizing critical errors.
Verified Tool Execution
All tool-triggered workflows and automations are now validated and reliable.
Full System Observability
Langfuse provides end-to-end visibility and metrics tracking, enabling faster debugging and optimization.
Scalable AI QA Infrastructure
Promptfoo + DeepEval + Langfuse created a continuous evaluation system, not a one-time QA effort.
“We didn’t just test the AI — we made it measurable, controllable, and ready to scale."
Engenious team
Looking Ahead
With a modern Playwright architecture and AI-powered workflows in place, the client is now positioned to scale automation faster than ever. The roadmap includes reaching 100% test coverage by February 2026, expanding AI-driven test generation, and further accelerating CI/CD execution. Forte Group continues to support the client in evolving their automation ecosystemб building a foundation that will only get faster, smarter, and more efficient over time.
Forte Group is redefining automation velocity with AI-driven Playwright transformation.
What can we scale together next?
Let’s redefine what quality can look like for your team.