Industrial IoT Release Quality Programme
Test strategy and automation architecture for a connected-plant platform where a bad release means stopped machines, not a broken page.
- Cypress
- TypeScript
- Postman
- REST APIs
- +2
Quality Engineering · GenAI Systems
Quality Engineering Leader specializing in GenAI Testing, RAG Validation, Automation Architecture, and Enterprise Software Quality.
Bengaluru, India · IST · UTC+5:30
A commerce graduate who taught himself to break software, then to build the systems that break it automatically — and now to measure the quality of systems that no longer give the same answer twice.
I did not start in engineering. I started in commerce, ledgers and reconciliation, where a single wrong number is not a cosmetic defect — it is a liability. That instinct never left. When I moved into computer applications and then into software testing, I brought the same question with me: what evidence do we actually have that this is correct?
For thirteen years that question has taken me through manual test design, automation frameworks, API and integration testing, and finally test leadership — owning release quality for Industrial IoT platforms and SEPA payment flows where a defect is measured in downtime or money.
Generative AI changed the shape of the question, not its substance. A deterministic assertion cannot validate a system that paraphrases. So I rebuilt my toolkit: retrieval quality metrics, groundedness and hallucination checks, LLM-as-judge evaluation harnesses, and regression suites for agent trajectories. Same discipline. New instruments.
Traditional testing asks “did it return the expected value?” AI quality engineering asks “is this answer supported, safe, and reproducible enough to ship?”
Periyar EVR College
Accounting rigour: reconciliation, audit trails, and an intolerance for unexplained variance.
Bishop Heber College
Formal computer science grounding — the deliberate pivot from finance into software.
Fhyzics Business Consultance
Test case design, functional, black-box and regression execution, defect lifecycle ownership, and production deployment support.
Capgemini
System integration, mobile and cross-browser testing at scale via Sauce Labs; Qualitia automation; UAT and production support on SharePoint, .NET and Java estates.
Trivium eSolutions
Owning test strategy, automation architecture and release governance across Industrial IoT and payments programmes. Mentoring engineers into a quality-first culture.
Applied AI Quality
Building and validating RAG pipelines, evaluation harnesses and agentic workflows — shipping AI systems that are trustworthy, not merely functional.
Not a list of tools. These are the problem spaces I own end to end — from strategy and architecture through to the evidence that ships with a release.
Strategy & Risk
Risk-based test strategy that maps coverage to business exposure instead of to screen count.
Frameworks & Scale
Maintainable automation architecture — deterministic, parallel-safe, and readable by the whole team.
Non-deterministic Systems
Test design for systems whose outputs vary — prompt contracts, tolerance bands, and semantic assertions.
Tool-Using Workflows
Validating multi-step agents on trajectory, not just final answer — tool calls, recovery, and termination.
Measurement
Evaluation harnesses with golden datasets, rubric-scored judges, and calibration against human review.
Retrieval Quality
Separating retrieval failure from generation failure — the single most common misdiagnosis in RAG.
Contracts & Integration
Contract-first API verification across REST services, message flows and third-party integrations.
Latency & Cost
Throughput, latency and — for AI systems — the token and cost budgets that decide viability.
Decision Discipline
Explicit entry and exit criteria so go/no-go is a data decision, never a hallway conversation.
Team Practice
ISTQB Agile Tester practice embedded in delivery: refinement, ATDD and continuous feedback.
Eight capabilities that make non-deterministic systems releasable. Each one exists because a specific failure mode kept reaching production.
Nominal, edge & adversarial cases, human-labelled
Prompt version, RAG config or agent policy
Answer + context, tool trajectory, tokens
calibrated vs. human
claim → source check
injection · PII · refusal
Per-class pass rate, delta vs. baseline
Gate blocks on regression against baseline
Failures don’t just fail the build — they become new regression cases added back into the golden dataset.
A prompt edit that improves one case silently regresses twenty others.
Version prompts as artefacts and run every candidate against a fixed scenario matrix — happy path, edge, adversarial, and multilingual — before promotion.
Each project is documented the way I would hand it over: the problem, the architecture, how it was tested, what it achieved, and what I would do differently.
A full-stack retrieval-augmented search agent that reads intent, retrieves semantically, and re-ranks with an LLM — built end-to-end and validated like production software.
PDF, DOCX & text sources
Section-aware, overlap window
Batched with idempotent upsert
Natural-language input
Same model — symmetry matters
Scoped before similarity
MongoDB · cosine similarity · ANN top-k = 50
Served on Groq · rubric-scored JSON
Every result traceable to its source chunk
Test strategy and automation architecture for a connected-plant platform where a bad release means stopped machines, not a broken page.
End-to-end validation of euro payment flows, where correctness is regulatory and a rejected file is a business incident.
A regression-testing framework for non-deterministic systems — golden datasets, rubric-scored judges, and calibration against human labels.
Quality work is only credible when it is countable. These are the numbers behind the narrative.
22 technologies I use in anger — organised by the problem they solve rather than by how impressive the logo looks.
Browser and end-to-end execution layers.
Cross-browser E2E, parallel & trace-first
Component and E2E suites in TypeScript
Legacy estate coverage and grid execution
Quality does not scale by hiring more testers. It scales when the whole team can reason about risk — and has the tooling to act on it.
Moving manual testers into automation: pairing on real suites, code review as teaching, and a deliberate path from execution to design.
Quality owned by delivery, not delegated to a gate at the end. Developers write tests, testers design strategy, and everyone reads the signal.
Coverage mapped to business risk and structured as a pyramid that actually holds — fast checks at the base, few and meaningful checks at the top.
Testing inside the sprint. Refinement produces testable acceptance criteria; automation lands with the story, not two sprints later.
Retrospectives with teeth: flake budgets, escaped-defect analysis, and cycle-time tracking that produce actions rather than sentiments.
Long-form pieces on making AI systems testable. Published as they are finished — the topics below are in the pipeline.
Exact-match assertions break the moment output becomes probabilistic. A working substitute: semantic assertions, tolerance bands, and per-class pass rates.
How to instrument retrieval and generation separately, and why Recall@k should be on the dashboard before anyone touches the prompt.
Decomposing answers into atomic claims, verifying each against context, and reporting unsupported-claim rate as a release metric.
Open to GenAI QA Specialist and GenAI Test Engineer roles, and to conversations about evaluation infrastructure, RAG validation and quality leadership.
Bengaluru, India · IST · UTC+5:30 · Open to remote & hybrid