This product was not featured by Product Hunt yet.
It will not be visible on their landing page and won't be ranked (cannot win product of the day regardless of upvotes).

Product upvotes vs the next 3

Waiting for data. Loading

Product comments vs the next 3

Waiting for data. Loading

Product upvote speed vs the next 3

Waiting for data. Loading

Product upvotes and comments

Waiting for data. Loading

Product vs the next 3

Loading

RepoGym

Validate the task before you benchmark the coding agent

Turn repository tasks into repeatable coding-agent evaluations. Define a task, check that doing nothing fails and the reference fix passes, then inspect scores and trajectories. Includes Python and JavaScript examples, composable graders, task mining and HTML reports. Free, MIT-licensed Python library. Start from the source checkout; local execution is for trusted code.

Top comment

I built RepoGym around a question that gets lost in agent benchmarks: can we trust the task itself? A broken baseline, a leaked test or a flaky grading command can make an agent score look meaningful when it is not. RepoGym puts the repository snapshot, test expectations, constraints and reference solution into a task you can inspect and rerun. Start with the included examples, validate the baseline and golden solution, then plug in your own agent. The library records trajectories and produces a standalone HTML report. It is free and MIT licensed; optional model providers have their own costs. I would love feedback from people building coding-agent evaluations: which validation failure has cost you the most time, and what would make this fit your repository? Source: https://github.com/shi1720/repogym Docs: https://shi1720.github.io/repogy... More of my work: https://shivamgupta.web.app/ Connect: https://www.linkedin.com/in/shiv...

About RepoGym on Product Hunt

Validate the task before you benchmark the coding agent

RepoGym was submitted on Product Hunt and earned 0 upvotes and 1 comments, placing #90 on the daily leaderboard. Turn repository tasks into repeatable coding-agent evaluations. Define a task, check that doing nothing fails and the reference fix passes, then inspect scores and trajectories. Includes Python and JavaScript examples, composable graders, task mining and HTML reports. Free, MIT-licensed Python library. Start from the source checkout; local execution is for trusted code.

On the analytics side, RepoGym competes within Developer Tools and GitHub — topics that collectively have 561.3k followers on Product Hunt. The dashboard above tracks how RepoGym performed against the three products that launched closest to it on the same day.

Who hunted RepoGym?

RepoGym was hunted by Shivam Gupta. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.

For a complete overview of RepoGym including community comment highlights and product details, visit the product overview.