Automate AI agent quality testing. Score correctness, detect hallucinations against ground truth, verify policy adherence, and benchmark RAG before bad responses reach real customers.
Hey Product Hunt! 👋
I’m thrilled to share QAgent with you today.
Here’s a dirty secret of modern AI engineering:
We write 50 unit tests for a 20-line backend function. But when we ship a non-deterministic AI agent handling real customers, our "testing pipeline" is typing 3 prompts into the OpenAI playground, seeing it respond politely, and hitting deploy.
The problem? Prompt regression is silent. You tweak one sentence in your system prompt to fix edge case A, and it silently breaks 3 other working customer flows without throwing a single runtime error.
Existing eval tools forced us to write 300 lines of custom Python scripts, pip install heavy libraries, and sift through terminal JSON dumps. We had to maintain an entire second Python codebase just to test our first one!
We built QAgent (https://qagent.in) to fix this for solo builders and agile teams:
⚡ 2-minute setup: Connect your agent via webhook or API endpoint. Zero SDK boilerplate.
🎯 8-dimension evaluation: Answer Quality, Factual Groundedness, Policy Adherence, Escalation Correctness, RAG Faithfulness, Contextual Relevancy, Context Recall, and Multi-turn Context Memory.
🛡️ Zero Python test scripts: Automated parallel test runs with visual pass/fail scorecards and root-cause failure breakdowns.
We’re offering 100 free evaluations every month (no credit card required) so every builder can stress-test their agent before shipping to customers.
I’d love to hear your thoughts, answer any questions, and get your most brutal feedback! 🚀
We just crossed 50 followers, 79 upvotes, and over 130 developers testing the platform since morning. Huge thank you to everyone who checked it out so far.
The biggest takeaway from the feedback and chats today is that almost everyone is struggling with the exact same failure mode: catching silent policy drift without getting buried in false alarms.
If anyone here is currently testing a customer-facing agent and wants help setting up custom evaluation webhooks, feel free to drop your use case below or ping me directly. Happy to help configure your ground-truth test sets today.
And remember, the free tier has 100 automated evaluations every month with zero SDK setup at qagent.in.
"Stop shipping on vibes" is uncomfortably accurate for where most of us are with agent QA right now - regression testing against ground truth instead of eyeballing transcripts is exactly the gap. Scoring policy adherence specifically is the part I'd actually pay for, since that's the failure mode that's hardest to catch by just reading a few sample outputs.
About QAgent on Product Hunt
“Automated QA for AI agents. Stop shipping on vibes.”
QAgent launched on Product Hunt on September 17th, 2026 and earned 80 upvotes and 6 comments, placing #20 on the daily leaderboard. Automate AI agent quality testing. Score correctness, detect hallucinations against ground truth, verify policy adherence, and benchmark RAG before bad responses reach real customers.
QAgent was featured in Developer Tools (519.8k followers), Artificial Intelligence (479.1k followers) and Bots (110.9k followers) on Product Hunt. Together, these topics include over 211.2k products, making this a competitive space to launch in.
Who hunted QAgent?
QAgent was hunted by Abhiram Reddy.K. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.
Want to see how QAgent stacked up against nearby launches in real time? Check out the live launch dashboard for upvote speed charts, proximity comparisons, and more analytics.