Skip to content
SA
Open to GenAI QA Specialist & GenAI Test Engineer roles

Quality Engineering · GenAI Systems

Engineering Confidencefor the AI Era

Quality Engineering Leader specializing in GenAI Testing, RAG Validation, Automation Architecture, and Enterprise Software Quality.

Years in Quality Engineering
13+
Years in Quality Engineering
Regulated & Industrial Domains
4
Regulated & Industrial Domains
ISTQB Certifications
2
ISTQB Certifications

Bengaluru, India · IST · UTC+5:30

01Who I Am

From Traditional QA to AI Quality Engineering

A commerce graduate who taught himself to break software, then to build the systems that break it automatically — and now to measure the quality of systems that no longer give the same answer twice.

I did not start in engineering. I started in commerce, ledgers and reconciliation, where a single wrong number is not a cosmetic defect — it is a liability. That instinct never left. When I moved into computer applications and then into software testing, I brought the same question with me: what evidence do we actually have that this is correct?

For thirteen years that question has taken me through manual test design, automation frameworks, API and integration testing, and finally test leadership — owning release quality for Industrial IoT platforms and SEPA payment flows where a defect is measured in downtime or money.

Generative AI changed the shape of the question, not its substance. A deterministic assertion cannot validate a system that paraphrases. So I rebuilt my toolkit: retrieval quality metrics, groundedness and hallucination checks, LLM-as-judge evaluation harnesses, and regression suites for agent trajectories. Same discipline. New instruments.

Traditional testing asks “did it return the expected value?” AI quality engineering asks “is this answer supported, safe, and reproducible enough to ship?”

Domains

  • Industrial IoT
  • Banking
  • SEPA Payments
  • Enterprise Applications

Certifications

  • ISTQB CTFL
  • ISTQB CTFL-AT

Languages

  • English
  • Tamil
  • Hindi
Career path
  1. 2004 — 2007Foundation

    B.Com, Commerce & Accounts

    Periyar EVR College

    Accounting rigour: reconciliation, audit trails, and an intolerance for unexplained variance.

  2. 2008 — 2011Transition

    Master of Computer Applications

    Bishop Heber College

    Formal computer science grounding — the deliberate pivot from finance into software.

  3. 2012 — 2014Manual Testing

    Software Engineer

    Fhyzics Business Consultance

    Test case design, functional, black-box and regression execution, defect lifecycle ownership, and production deployment support.

  4. 2014 — 2017Automation

    Consultant — Test Engineer

    Capgemini

    System integration, mobile and cross-browser testing at scale via Sauce Labs; Qualitia automation; UAT and production support on SharePoint, .NET and Java estates.

  5. 2017 — PresentQuality Leadership

    Test Lead

    Trivium eSolutions

    Owning test strategy, automation architecture and release governance across Industrial IoT and payments programmes. Mentoring engineers into a quality-first culture.

  6. 2024 — PresentGenAI Engineering

    GenAI / LLM Test Engineer

    Applied AI Quality

    Building and validating RAG pipelines, evaluation harnesses and agentic workflows — shipping AI systems that are trustworthy, not merely functional.

02Technical Expertise

Ten engineering domains, one discipline

Not a list of tools. These are the problem spaces I own end to end — from strategy and architecture through to the evidence that ships with a release.

01

Strategy & Risk

Quality Engineering

Risk-based test strategy that maps coverage to business exposure instead of to screen count.

  • Risk-based test strategy
  • Test architecture & design
  • Shift-left quality gates
02

Frameworks & Scale

Automation Engineering

Maintainable automation architecture — deterministic, parallel-safe, and readable by the whole team.

  • Playwright & Cypress frameworks
  • Page-object & fixture design
  • Parallel execution & sharding
03

Non-deterministic Systems

GenAI Testing

Test design for systems whose outputs vary — prompt contracts, tolerance bands, and semantic assertions.

  • Prompt regression suites
  • Semantic equivalence assertions
  • Adversarial & jailbreak probes
04

Tool-Using Workflows

Agent Testing

Validating multi-step agents on trajectory, not just final answer — tool calls, recovery, and termination.

  • Trajectory & step-order assertions
  • Tool-call contract validation
  • Loop, budget & timeout guards
05

Measurement

LLM Evaluation

Evaluation harnesses with golden datasets, rubric-scored judges, and calibration against human review.

  • Golden dataset curation
  • LLM-as-judge rubrics
  • Judge calibration & agreement
06

Retrieval Quality

RAG Validation

Separating retrieval failure from generation failure — the single most common misdiagnosis in RAG.

  • Recall@k / MRR / nDCG tracking
  • Chunking & embedding ablations
  • Groundedness & citation checks
07

Contracts & Integration

API Testing

Contract-first API verification across REST services, message flows and third-party integrations.

  • Postman & REST Assured suites
  • Schema & contract testing
  • Auth, RBAC & negative paths
08

Latency & Cost

Performance Testing

Throughput, latency and — for AI systems — the token and cost budgets that decide viability.

  • Load & soak profiling
  • Vector search benchmarking
  • Token & inference cost analysis
09

Decision Discipline

Release Governance

Explicit entry and exit criteria so go/no-go is a data decision, never a hallway conversation.

  • Entry / exit criteria design
  • Release readiness reporting
  • UAT & production-support flow
10

Team Practice

Agile QA

ISTQB Agile Tester practice embedded in delivery: refinement, ATDD and continuous feedback.

  • Story refinement & ACs
  • ATDD / BDD collaboration
  • In-sprint automation
03GenAI & AI Testing Lab

Quality engineering for systems that never answer twice the same way

Eight capabilities that make non-deterministic systems releasable. Each one exists because a specific failure mode kept reaching production.

Fig. 01 — LLM evaluation harnessdataset → target → judges → gate
Capability index
EVAL-01

Prompt Testing

Failure mode

A prompt edit that improves one case silently regresses twenty others.

How I test it

Version prompts as artefacts and run every candidate against a fixed scenario matrix — happy path, edge, adversarial, and multilingual — before promotion.

Signals tracked

  • Pass rate by scenario class
  • Format conformance
  • Regression delta vs. baseline
04Projects

Case studies, not screenshots

Each project is documented the way I would hand it over: the problem, the architecture, how it was tested, what it achieved, and what I would do differently.

FeaturedGenAI · Retrieval Systems2025

RAG Resume Search Agent

A full-stack retrieval-augmented search agent that reads intent, retrieves semantically, and re-ranks with an LLM — built end-to-end and validated like production software.

  • Node.js
  • TypeScript
  • MongoDB Atlas Vector Search
  • Mistral Embeddings
  • Groq
  • Llama
Semantic
Intent-level matching replaces boolean keyword strings
2-stage
Wide ANN recall, then precision re-ranking
100%
Results returned with traceable source citations
Sub-second
Retrieval latency before re-rank on the candidate set
Source
Fig. 02 — RAG Resume Search Agenttwo-stage retrieval · recall then precision
Retrieval strategy
Wide ANN recall, then LLM precision re-rank
Embedding symmetry
Identical model for documents and queries
Trust mechanism
Every result traceable to its source chunk
Industrial IoT · Test Leadership

Industrial IoT Release Quality Programme

Test strategy and automation architecture for a connected-plant platform where a bad release means stopped machines, not a broken page.

  • Cypress
  • TypeScript
  • Postman
  • REST APIs
  • +2
2018 — PresentCase study
Banking · Payments

SEPA Payments Assurance

End-to-end validation of euro payment flows, where correctness is regulatory and a rejected file is a business incident.

  • Postman
  • SQL
  • ISO 20022 / XML
  • HP ALM
  • +2
2015 — 2019Case study
GenAI · Evaluation Infrastructure

LLM Evaluation Harness

A regression-testing framework for non-deterministic systems — golden datasets, rubric-scored judges, and calibration against human labels.

  • TypeScript
  • Node.js
  • Groq / Llama
  • Mistral
  • +2
2025Case study
05Engineering Impact

Thirteen years, measured

Quality work is only credible when it is countable. These are the numbers behind the narrative.

0+
Years of Experience
Software quality since 2012
0+
Projects Delivered
Across four industry domains
0
Automation Frameworks
Designed, built and maintained
0+
Engineers Mentored
Manual testers to SDETs
0+
Enterprise Releases
Governed to exit criteria
0+
AI Experiments
RAG, eval and agent workflows
06Tech Stack

The tooling map, grouped by job

22 technologies I use in anger — organised by the problem they solve rather than by how impressive the logo looks.

Automation

layer / automation

Browser and end-to-end execution layers.

  • Playwright

    Cross-browser E2E, parallel & trace-first

  • Cypress

    Component and E2E suites in TypeScript

  • Selenium

    Legacy estate coverage and grid execution

PlaywrightCypressSeleniumTypeScriptNode.jsJavaPythonPostmanREST AssuredMistralLlamaGroqRAGVector SearchMongoDBSQLAzure DevOpsGitHubJenkinsAgile / ScrumJIRAHP ALMPlaywrightCypressSeleniumTypeScriptNode.jsJavaPythonPostmanREST AssuredMistralLlamaGroqRAGVector SearchMongoDBSQLAzure DevOpsGitHubJenkinsAgile / ScrumJIRAHP ALM
07Leadership

Building Quality-First Engineering Teams

Quality does not scale by hiring more testers. It scales when the whole team can reason about risk — and has the tooling to act on it.

01

Mentoring

Moving manual testers into automation: pairing on real suites, code review as teaching, and a deliberate path from execution to design.

  • Pairing over lecturing
  • Manual → SDET pathways
02

Quality Culture

Quality owned by delivery, not delegated to a gate at the end. Developers write tests, testers design strategy, and everyone reads the signal.

  • Shared ownership
  • Blameless defect analysis
03

Testing Strategy

Coverage mapped to business risk and structured as a pyramid that actually holds — fast checks at the base, few and meaningful checks at the top.

  • Risk-weighted coverage
  • Layer-appropriate tests
04

Agile Delivery

Testing inside the sprint. Refinement produces testable acceptance criteria; automation lands with the story, not two sprints later.

  • Testable ACs at refinement
  • In-sprint automation
05

Continuous Improvement

Retrospectives with teeth: flake budgets, escaped-defect analysis, and cycle-time tracking that produce actions rather than sentiments.

  • Flake budgets
  • Escaped-defect RCA

Operating principles

  • A test that nobody trusts is worse than no test at all.
  • Escaped defects are a process signal, not an individual failure.
  • If the release decision cannot be explained with data, it was not a decision.
  • Automation exists to buy attention for the testing only humans can do.
08Writing & Insights

Notes from the evaluation lab

Long-form pieces on making AI systems testable. Published as they are finished — the topics below are in the pipeline.

Get notified
Drafting9 min

Testing What Does Not Repeat: A Practical Model for GenAI QA

Exact-match assertions break the moment output becomes probabilistic. A working substitute: semantic assertions, tolerance bands, and per-class pass rates.

GenAI Testing
Drafting11 min

Your RAG Bug Is Probably a Retrieval Bug

How to instrument retrieval and generation separately, and why Recall@k should be on the dashboard before anyone touches the prompt.

RAG Evaluation
Planned10 min

Claim-Level Hallucination Detection Without a Human in the Loop

Decomposing answers into atomic claims, verifying each against context, and reporting unsupported-claim rate as a release metric.

LLM Hallucination Testing
Contact

Let’s Build Trustworthy AI and Software Systems

Open to GenAI QA Specialist and GenAI Test Engineer roles, and to conversations about evaluation infrastructure, RAG validation and quality leadership.

Bengaluru, India · IST · UTC+5:30 · Open to remote & hybrid