EnGenious
Phoenix AI

Bringing Modern, AI-Driven Quality to Legacy Systems

 

At-a-Glance

 

  • Platforms: LangGraph Multi-Agent System, Web & Mobile Apps, DeepEval, Promptfoo, Langfuse

  • Core Features: AI routing engine, dynamic agent handoffs, tool-driven automation, metrics tracking & evaluation

  • Smooth Execution: Synthetic dataset validation, trace-level observability, scalable evaluation pipelines

  • Delivery Impact: 5,000+ scenarios validated, 2× faster feature delivery driven by stable, measurable AI quality gates

  • Differentiator: Full-stack AI evaluation system combining DeepEval + Promptfoo + Langfuse with metrics-driven validation

 

Aba243c7 a0af 4177 bc54 9bc5a7965850

 

Background

 
Phoenix AI is building a multi-agent AI platform designed to handle complex, real-world user interactions through dynamic agent orchestration.

Instead of a single chatbot, the system relies on LangGraph-based routing, where a primary agent continuously decides which specialist should respond.

This unlocks powerful capabilities — but also introduces a critical risk:

Without evaluation, the system can silently fail — routing users incorrectly, breaking context, and producing inconsistent outcomes.

We transform this into a controlled, testable, production-ready system.

 

Goals & Objectives

 
The objective was to make AI behavior predictable under pressure, measurable, and continuously improvable.

These included:

  • Ensure Routing Accuracy: Every input maps to the correct agent
  • Control Agent Switching: Eliminate unnecessary or incorrect handoffs
  • Validate End-to-End Flows: Ensure reliable execution across workflows and tools
  • Test at Scale: Simulate thousands of real-world scenarios
  • Establish Quality Gates: Introduce measurable validation before every release
  • Track Key Metrics: Monitor routing accuracy, failures, and system behavior over time
     

 

Forte img 2

 

Challenges

 
Multi-agent AI systems fail differently than traditional software — and often invisibly.

 

1. Invisible Failures
Incorrect routing or outputs don’t crash systems — they degrade trust silently.

2. Non-Deterministic Behavior
The same input produces different outputs, requiring semantic and statistical validation.

3. Routing Is the Core Risk
The system must decide which agent responds — making routing accuracy critical.

4. Fragile Conversation Flow
Uncontrolled switching breaks continuity and reduces usability.

5. Tool & Workflow Reliability
AI doesn’t just respond — it triggers tools and workflows that must execute correctly.

6. No Unified Testing Stack
Validation requires combining multiple tools with clearly defined roles.

 

Partnership & Process

We built a layered AI evaluation system, where each tool has a clear responsibility and measurable output.

 

Together, we:

1. DeepEval — Synthetic Dataset & Semantic Evaluation

  • Utilise DeepEval to generate and validate 5,000+ synthetic scenarios:
  • Realistic user journeys
  • Edge cases and ambiguous inputs
  • Long conversational flows
  • Rare and high-risk situations
  • DeepEval enabled semantic scoring and large-scale scenario validation.
     

2. Promptfoo — Regression & Behavior Testing

  • Automated prompt validation
  • Defined expected behaviors and scoring rules
  • Regression testing across updates
  • Output comparison across runs
  • Ensuring no regressions reach production.
     

3. Langfuse — Observability & Metrics Tracking

  • Full trace-level visibility
  • Agent routing inspection
  • Prompt-response debugging
  • Metrics tracking (routing accuracy, failure rates, latency, flow success)
  • Turning AI into a data-driven system.
     

4. Tool & Workflow Evaluation

  • Tool invocation accuracy
  • Cross-system automation reliability
  • Input/output consistency
  • Failure handling across workflows

 



 

 

Forte img 4

Explore Our QA Services 

 

Outcomes & Impact

 
This engagement didn’t just improve quality — it made AI measurable, faster to evolve, and safe to scale.

 

2× Faster Test Development
Structured evaluation enabled 2× faster delivery of new agentic features (200%+ increase in velocity).

 

Measurable AI Quality Gates
Every release now passes through quantifiable validation metrics, reducing production risk.

 

5,000+ Scenarios Validated
DeepEval-powered datasets ensured broad coverage and early detection of failures.

 

Improved Routing Accuracy
Agent selection became consistent and reliable — minimizing critical errors.

 

Verified Tool Execution
All tool-triggered workflows and automations are now validated and reliable.

 

Full System Observability
Langfuse provides end-to-end visibility and metrics tracking, enabling faster debugging and optimization.
 

Scalable AI QA Infrastructure
Promptfoo + DeepEval + Langfuse created a continuous evaluation system, not a one-time QA effort.
 

“We didn’t just test the AI — we made it measurable, controllable, and ready to scale."

Engenious team

Looking Ahead

 

With a modern Playwright architecture and AI-powered workflows in place, the client is now positioned to scale automation faster than ever. The roadmap includes reaching 100% test coverage by February 2026, expanding AI-driven test generation, and further accelerating CI/CD execution. Forte Group continues to support the client in evolving their automation ecosystemб building a foundation that will only get faster, smarter, and more efficient over time.

 

 

Forte Group is redefining automation velocity with AI-driven Playwright transformation.
What can we scale together next?

 
Let’s redefine what quality can look like for your team.

 

Start a Conversation