Non-Determinism
Same prompt can yield different responses. Traditional assert-equals testing fails. Requires semantic similarity, rubric scoring, and statistical evaluation over multiple runs.
If you've spent years testing traditional software, AI agents will humble you quickly. The assumptions that make classic QA work (same input, same output, assert-equals) simply don't hold here. The same prompt can return a different answer on every run, agents decide for themselves which tools to call, and multi-agent workflows produce emergent behavior that no single unit test can predict.
Same prompt can yield different responses. Traditional assert-equals testing fails. Requires semantic similarity, rubric scoring, and statistical evaluation over multiple runs.
Agents autonomously select tools, APIs, and parameters. Wrong tool selection or malformed parameters cause silent failures that propagate downstream.
Agent collaboration creates behavior that no single agent exhibits alone. Orchestration bugs, tool failures, and misalignment are invisible in unit tests.
Token limits, stale vector DB entries, and context window mismanagement errors degrade quality gradually, not suddenly. Hard to catch without specialized tools.
Prompt injection, PII leakage, hallucination, and guardrail bypasses represent new attack vectors with no equivalent in traditional software.
None of these gaps closes on its own, and none of them yields to traditional QA. What works is a purpose-built quality engineering practice: AI-native evaluation techniques paired with enterprise-grade process rigor. That's exactly what the Innocito AI Testing Framework is built to give you.
We organize AI testing into seven progressive phases, each building on the one before it. The logic is simple: catch every defect at the earliest, and cheapest, stage where it can be caught.
| PHASE | WHAT IT VALIDATES | KEY TOOLS |
|---|---|---|
| 1. Agent Unit & Prompt | Individual agent correctness, prompt stability, schema compliance | Pytest, DeepEval, LangSmith |
| 2. Workflow & Integration | Multi-agent handoffs, orchestration, planner-executor loops | LangSmith, Playwright, custom harness |
| 3. Context & Memory | RAG retrieval, context window, memory consistency, long conversations | Ragas, TruLens, MTEB |
| 4. Adversarial & Fuzzing | Prompt injection, malformed inputs, edge cases, contradictions | Garak, PyRIT, Promptfoo |
| 5. Resilience & Chaos | Failure recovery, retries, circuit breakers, degraded operation | Toxiproxy, Chaos Toolkit |
| 6. Performance & Load | Latency, throughput, scalability, soak stability | Locust, k6, OpenTelemetry |
| 7. Guardrails & Safety | PII, toxicity, hallucination, policy compliance, output scope | Presidio, Garak, Perspective API |
The sections that follow walk through each phase in detail: step-by-step procedures, concrete examples, tool configurations, and checklists you can put to work the same day.
Thresholds like a 95% task success rate, a sub-2% hallucination rate, and P95 latency under 8 seconds aren't universal constants. They're the starting points we calibrate to across AI testing engagements, defensible defaults, not guesses. Tune them to your own domain’s risk tolerance; treat them as a baseline to argue with, not a certification to pass.
Validate individual agents in isolation before they participate in any workflow.
Document each agent's input schema, output schema, system prompt, and tool access list. Maintain this catalogue in a shared wiki or API spec (OpenAPI / JSON Schema).
Create a minimum of 50 labeled input/expected-output pairs per agent. Include: happy path (60%), edge cases (20%), adversarial inputs (10%), boundary conditions (10%). Tool: LangSmith Datasets or Argilla for collaborative dataset curation.
For each agent, define evaluation criteria: structured output uses exact match or JSON schema validation; free text uses semantic similarity (cosine ≥ 0.85) plus ROUGE-L ≥ 0.70; tool selection asserts the correct tool and valid parameters for each test case. Tools: DeepEval (metrics library), Pytest (runner), LangSmith (tracing and evaluation).
Execute the evaluation dataset against the current prompt version. Store results with the prompt version hash. Any prompt change that triggers re-evaluation; alert if accuracy drops more than 2% versus baseline. CI integration: add as a GitHub Actions / Jenkins step that blocks merge on failure.
Because LLM outputs are non-deterministic, run each test case 3 to 5 times. Use majority vote or mean score to determine pass/fail. Flag cases with high variance (more than 20% spread) for prompt refinement.
Agent: SummarizerV2 | Prompt version: v2.3 | Dataset: 80 labeled articles
What we see in practice: teams build the happy-path 60% of the evaluation dataset first, ship the agent once it's green, and never circle back to the adversarial and boundary slices. Those are exactly the cases that don't show up in a demo, and do show up in production a few months later.
Validate agent collaboration, handoffs, and end-to-end pipeline correctness.
User request: "Produce a 500-word competitive analysis of electric vehicle battery manufacturers, citing 3 sources."
What we see in practice: the handoff contract test is the first thing cut under deadline pressure, and it's almost always the root cause when a pipeline silently returns garbage instead of failing loudly. Writing it up front costs a day. Debugging its absence in production costs a lot more.
Ensure agents retrieve the right information and retain context across interactions
What we see in practice: the handoff contract test is the first thing cut under deadline pressure, and it's almost always the root cause when a pipeline silently returns garbage instead of failing loudly. Writing it up front costs a day. Debugging its absence in production costs a lot more.
Probe the system with unexpected, malformed, and malicious inputs to find hidden failure modes
What we see in practice: automated scanners catch the attacks everyone already knows about. What actually breaks systems in production is more ordinary: a support ticket pasted into a chat window that happens to look like a system prompt. Fuzzing needs to cover that kind of accidental chaos, not just deliberate red-team scripts.
Prove the system degrades gracefully when components slow down, fail, or disappear entirely
What we see in practice: teams that skip chaos testing usually have a retry policy. They just don't know if it's a good one until the first real outage, when retries without backoff turn a slow dependency into a self-inflicted traffic spike.
Prove the system holds its latency, throughput, and stability targets under real-world traffic.
| SCENARIO | DESCRIPTION | DURATION | PASS CRITERION |
|---|---|---|---|
| Baseline | 10% peak users | 30 min | P95 < 5s, errors < 0.1% |
| Normal Peak | Expected peak traffic | 60 min | P95 < 8s, errors < 0.5% |
| Stress | 150% of peak | 30 min | P95 < 15s, errors < 2% |
| Spike | 10x traffic in 30 sec | 10 min | Recovers within 2 min |
| Soak | 60% peak sustained | 8–24 hrs | No drift / > 20% from baseline |
| Breakpoint | Ramp until failure | Until break | Identify saturation point |
What we see in practice: teams that skip chaos testing usually have a retry policy. They just don't know if it's a good one until the first real outage, when retries without backoff turn a slow dependency into a self-inflicted traffic spike.
Ensure all outputs comply with safety, privacy, and organizational policy standards
What we see in practice: guardrails validated once before launch create a false sense of safety. Every model update, prompt change, and new tool integration is a fresh chance for a previously blocked pattern to get through. The teams that stay safe treat this phase as continuous, not a one-time gate.
Not sure which test to run? Work through the questions below: they’ll route you to the right phase in under a minute.
Run Phase 1: Agent Unit & Prompt Testing. Re-run the evaluation dataset.
Run Phase 2: Workflow & Integration Testing. Validate handoff contracts and macro outcomes.
Run Phase 3: Context & Memory Validation. Benchmark retrieval accuracy and staleness.
Run Phase 4 + Phase 7: full adversarial fuzzing and guardrail validation.
Run Phase 5 + Phase 6: resilience, load, and soak testing. If none of the above apply, run the full Phase 1–7 suite: a major change has been detected
| SYMPTOM | LIKELY ROOT CAUSE | FIRST TEST TO RUN | ESCALATION |
|---|---|---|---|
| Wrong output format | Prompt regression or schema change | Phase 1: prompt regression suite | Dev → prompt review |
| Incorrect facts in output | Hallucination or stale memory | Phase 3: retrieval accuracy + Phase 7: faithfulness | QA → red team |
| Pipeline produces no output | Handoff failure or agent crash | Phase 2: trace-based integration test | Dev → trace analysis |
| Slow response times | LLM latency or queue backlog | Phase 6: latency breakdown via tracing | Dev → infra team |
| Offensive or harmful output | Guardrail bypass | Phase 7: toxicity + injection test | QA → immediate hotfix |
| Intermittent failures | Resource exhaustion or race condition | Phase 5: chaos test + Phase 6: soak | Dev → memory profiling |
These are the same checklists our teams run on real engagements. Copy them, customize them, or drop them straight into your project management tool.
A tool only earns its place once it's running. These condensed guides take each core tool in the framework from zero to first useful result, fast.
LLM Evaluation & Tracing
End-to-end LLM tracing, dataset management, automated evaluation pipeline
pip install langsmith && export LANGCHAIN_API_KEY=<your-key>
Trace every LLM call with latency/token stats; create evaluation datasets; run evaluators (exact match, semantic similarity, custom); CI/CD integration via CLI
Phase 1 (prompt regression), Phase 2 (workflow tracing), Phase 3 (memory evaluation)
langsmith evaluate --dataset my_eval_set --evaluator semantic_similarity --model gpt-4
A strong default if you're already in the LangChain ecosystem. Teams on a different stack sometimes find its tracing model assumes more about your architecture than they'd like.
Llm Vulnerability scanner
Automated adversarial testing: prompt injection, jailbreak, data leakage, hallucination probing
pip install garak
200+ built-in attack probes; supports OpenAI, Anthropic, Hugging Face, custom endpoints; generates HTML reports
Phase 4 (adversarial fuzzing), Phase 7 (guardrail validation)
garak --model_type openai --model_name gpt-4--probes all
An excellent first pass, but its probe library goes stale as fast as jailbreak techniques evolve. A clean Garak run is a floor, not a ceiling.
RAG Evaluation Framework
Evaluate retrieval-augmented generation quality: faithfulness, answer relevancy, context precision/recall
pip install ragas
Faithfulness (is the answer grounded in context?), answer relevancy (does it address the question?), context precision/recall (are the right docs retrieved?)
Phase 3 (context and memory validation)
from ragas import evaluate; results = evaluate(dataset, metrics=[faithfulness, answer_relevancy, context_precision])
The metrics are only as good as your groundtruth dataset. Point it at synthetic data instead of a real evaluation set and you'll get numbers that look great and mean very little.
Load Testing
Python-native load testing with scriptable user behavior; distributed mode for high concurrency
pip install locust
Define user workflows as Python classes; real-time web dashboard; CSV export; distributed workers
Phase 6 (load, stress, spike, soak testing)
locust -f agent_load.py -- host=https://api.example.com --users=100 --spawn-rate=10
The easiest way to load test in Python, but its default user-behavior model doesn't capture bursty, cascading agent-to-agent calls well. Plan to hand-roll a task set that reflects real agent chains, not single request/response pairs.
Network Fault Injection
Simulate network faults between services: latency, bandwidth throttle, connection drops, timeouts
docker run -p 8474:8474 shopify/toxiproxy(or: brew install toxiproxy)
Create proxies per dependency; add/remove toxics (latency, bandwidth, timeout) via API or CLI; zero code changes required
Phase 5 (resilience and chaos testing)
toxiproxy-cli create llm_api -l 0.0.0.0:5555-u api.openai.com:443 && toxiproxy-cli toxic add llm_api -t latency -a latency=5000
It does one job and stays out of the way. The real value isn't the tool, it's that most teams don't write resilience assertions until they have a way to actually inject failure. Once they do, they rarely go back to testing without it.
PII Detection
Detect and anonymize PII (names, SSNs, emails, credit cards, medical records) in text
pip install presidio-analyzer presidioanonymizer
40+ built-in PII recognizers; NLP + regex hybrid detection; customizable entity types; anonymization operators (redact, hash, replace)
Phase 7 (guardrail validation, PII leakage testing)
from presidio_analyzer import AnalyzerEngine; analyzer = AnalyzerEngine(); results = analyzer.analyze(text=agent_output, language='en')
Strong at the entity types it knows out of the box. It won't catch domain-specific sensitive data, internal project codenames, proprietary identifiers, without extending its recognizers first. Don't treat it as a complete PII gate on day one.
Every organization starts this journey from a different place, so we’ve built engagement models to match: whether you need a quick assessment, a full framework build-out, or an embedded team that owns AI quality end to end.
A 2 to 4 week evaluation of your current AI testing maturity. Deliverable: gap analysis, prioritized roadmap, tool recommendations, and estimated effort.
End-to-end implementation of the seven-phase testing framework. Includes tool setup, CI/CD integration, evaluation dataset creation, and knowledge transfer.
A dedicated Innocito QE team embedded in your sprints. Continuous testing, redteaming, load testing, and quality reporting. Monthly cadence.
Periodic adversarial testing campaigns. Our security-focused AI testers probe your system with the latest attack techniques. Quarterly or on demand.
Hands-on workshops for your QA, Dev, and BA teams. Covers AI testing fundamentals, tool proficiency, and framework adoption. One to three day sessions.
Tell us where your AI testing stands today, and we'll map the fastest path to where it needs to be. Book a consultation with our AI Quality Engineering team.