Observed
Five sessions did not complete. Half of all actions were scrolling. Every session reached for chat.
Behavioral regression testing for AI-built web apps
JinuGen runs the same six simulated testers through one critical preview flow in a real browser. See where they loop, stall or abandon—then rerun the same panel after your fix to see what changed.
One flow · six simulated sessions · evidence-backed findings · one rerun
Actual run · customer anonymized
The goal was to find the ages, schedule and price for live online classes. No page returned an error and no control was dead. Yet five simulated sessions did not complete, and the sixth reached the end of its steps without the answer either — the page contained no class times and no prices.
5 / 6
sessions never got through
50%
of all 100 actions were scrolling
6 / 6
sessions reached for chat
0
class times or prices on the page
STORED ACTION TRACE
Jordan · 13 actions · 03:02
“Scroll down further past this marketing fluff section to find actual schedule times, pricing tiers, and sign-up”
Model rationale recorded before the action. This is an animation of stored telemetry, not browser video.
Observed
Five sessions did not complete. Half of all actions were scrolling. Every session reached for chat.
Interpretation
The sessions kept searching beyond the primary page because the journey’s decision information was absent.
Limit
One page, one goal and one run configuration. This is not an estimate of real-customer conversion.
Keep your existing tests
FUNCTIONAL TESTS
JINUGEN
CUSTOMER RESEARCH
JinuGen replaces neither. It covers the fast feedback loop between a working build and the next customer conversation.
One baseline · one comparable rerun
Give us a staging URL and the first-use journey that matters. No code integration is required for the pilot.
Six sessions independently work the goal in real Chromium. We capture actions, screenshots, outcomes and stated rationale.
Findings are tied to observable behavior, reproduction steps and affected-session counts.
We reuse the goal, panel, device and configuration, then report what is new, still present, fixed or unverified.
The profiles and run configuration stay fixed. The model is stochastic, so individual clicks are not guaranteed to repeat. If the rerun does not exercise the relevant page, the finding is marked unverified—not fixed.
Founding design-partner pilot · $500
For teams shipping web interfaces faster than they can manually pressure-test every first-use journey.
Limited to the first five qualified teams.
GOOD FIT
A browser-based preview you control, test data and one journey with a clear end state.
NOT A FIT
Live customer data, native mobile, destructive payments, broad audits or targets you do not control.
The fee covers execution and analysis, not a promised defect count. If our tooling compromises the run, we rerun it at our cost or refund the pilot.
Founding design-partner pilot · $500
We reply personally within one business day. Nothing runs until scope, authorization and payment are confirmed with you.
First-Use Friday · free
Each week we pick one authorized preview flow, run the panel against it, and publish the result as an anonymized case study. You get the full private report either way. We get evidence we are allowed to show.
WHAT YOU GET
The same six sessions and the same evidence-backed findings as the paid pilot, on one flow. Reviewed by hand before you see it.
WHAT WE ASK
A preview or staging target you own, synthetic test data, and written permission to publish a sanitized version. Nothing is published without your approval of the exact wording.
No guaranteed selection and no promised date — we take one a week and say no to the rest. No live customer data, no destructive flows, no targets you do not control. The paid pilot below is private, scheduled and includes the rerun; this lane is neither.
Private by default
Founding pilots use authorized preview environments and test data. Credentials are handled separately, never through the public form.
Questions buyers ask
No. They are model-generated sessions with fixed behavioral parameters. They do not predict population behavior or replace customer research.
Keep Playwright. It verifies behavior you already knew to assert. JinuGen explores how a simulated first-time session tries to achieve a goal and surfaces repeated breakdowns for investigation.
Not for founding pilots. We test authorized preview or staging environments using test data.
We say so plainly. You still receive completion outcomes and evidence; we do not manufacture a quota of findings.
No. The pilot is manually scoped and delivered. You provide a URL, goal and any temporary test access agreed privately.
Request one of five founding design-partner pilots. We reply within one business day and run nothing until scope, authorization and payment are confirmed.
Request the founding pilot — $500