Behavioral regression testing for AI-built web apps

Catch the first-use failures your test suite calls green.

JinuGen runs the same six simulated testers through one critical preview flow in a real browser. See where they loop, stall or abandon—then rerun the same panel after your fix to see what changed.

One flow · six simulated sessions · evidence-backed findings · one rerun

What this signal is: model-generated sessions with fixed behavioral parameters—not real users or a statistically representative sample. JinuGen complements functional tests and customer research.

Actual run · customer anonymized

Nothing was broken. Nobody got the answer.

The goal was to find the ages, schedule and price for live online classes. No page returned an error and no control was dead. Yet five simulated sessions did not complete, and the sixth reached the end of its steps without the answer either — the page contained no class times and no prices.

5 / 6

sessions never got through

50%

of all 100 actions were scrolling

6 / 6

sessions reached for chat

0

class times or prices on the page

STORED ACTION TRACE

Jordan · 13 actions · 03:02

ACTION 04 · +00:42scroll
Scroll down further past this marketing fluff section to find actual schedule times, pricing tiers, and sign-up

Model rationale recorded before the action. This is an animation of stored telemetry, not browser video.

Observed

Five sessions did not complete. Half of all actions were scrolling. Every session reached for chat.

Interpretation

The sessions kept searching beyond the primary page because the journey’s decision information was absent.

Limit

One page, one goal and one run configuration. This is not an estimate of real-customer conversion.

Open the complete redacted action trace →

Keep your existing tests

Three questions. Three different tools.

FUNCTIONAL TESTS

Did the behavior we knew to assert still work?

CUSTOMER RESEARCH

What do actual customers need, understand and value?

JinuGen replaces neither. It covers the fast feedback loop between a working build and the next customer conversation.

One baseline · one comparable rerun

From preview link to verified change.

  1. 01

    Send one preview and one goal

    Give us a staging URL and the first-use journey that matters. No code integration is required for the pilot.

  2. 02

    Run the fixed panel

    Six sessions independently work the goal in real Chromium. We capture actions, screenshots, outcomes and stated rationale.

  3. 03

    Review evidence, not AI prose

    Findings are tied to observable behavior, reproduction steps and affected-session counts.

  4. 04

    Fix and rerun

    We reuse the goal, panel, device and configuration, then report what is new, still present, fixed or unverified.

The profiles and run configuration stay fixed. The model is stochastic, so individual clicks are not guaranteed to repeat. If the rerun does not exercise the relevant page, the finding is marked unverified—not fixed.

Founding design-partner pilot · $500

One critical flow. Baseline, fix, rerun.

For teams shipping web interfaces faster than they can manually pressure-test every first-use journey.

Limited to the first five qualified teams.

  • Six simulated sessions against one preview or staging flow
  • One desktop or mobile run configuration
  • Evidence-backed findings with screenshots and reproduction steps
  • A concise prioritized report and 30-minute founder review
  • One comparable rerun within 14 days

GOOD FIT

A browser-based preview you control, test data and one journey with a clear end state.

NOT A FIT

Live customer data, native mobile, destructive payments, broad audits or targets you do not control.

The fee covers execution and analysis, not a promised defect count. If our tooling compromises the run, we rerun it at our cost or refund the pilot.

Founding design-partner pilot · $500

Tell us where a first-time user must get through.

We reply personally within one business day. Nothing runs until scope, authorization and payment are confirmed with you.

Do not include passwords, signed URLs, tokens or other secrets.

Submitting does not start a run. No card, automated production scan or sales sequence.

First-Use Friday · free

One preview a week, tested in public.

Each week we pick one authorized preview flow, run the panel against it, and publish the result as an anonymized case study. You get the full private report either way. We get evidence we are allowed to show.

WHAT YOU GET

The same six sessions and the same evidence-backed findings as the paid pilot, on one flow. Reviewed by hand before you see it.

WHAT WE ASK

A preview or staging target you own, synthetic test data, and written permission to publish a sanitized version. Nothing is published without your approval of the exact wording.

No guaranteed selection and no promised date — we take one a week and say no to the rest. No live customer data, no destructive flows, no targets you do not control. The paid pilot below is private, scheduled and includes the rerun; this lane is neither.

Put a flow forward

Private by default

A browser test can see sensitive things. We treat that as part of the product.

Founding pilots use authorized preview environments and test data. Credentials are handled separately, never through the public form.

  • Submitting this form never starts a run.
  • Raw visual artifacts default to deletion within 14 days of delivery.
  • Customer content is not used to train or benchmark JinuGen without written opt-in.
  • Observed behavior is reported separately from model interpretation.
  • A finding is called fixed only after the relevant page is exercised again.

Questions buyers ask

Plain answers.

Are these real users?

No. They are model-generated sessions with fixed behavioral parameters. They do not predict population behavior or replace customer research.

Why not just use Playwright?

Keep Playwright. It verifies behavior you already knew to assert. JinuGen explores how a simulated first-time session tries to achieve a goal and surfaces repeated breakdowns for investigation.

Can you test production?

Not for founding pilots. We test authorized preview or staging environments using test data.

What if JinuGen finds nothing?

We say so plainly. You still receive completion outcomes and evidence; we do not manufacture a quota of findings.

Do I need to install anything?

No. The pilot is manually scoped and delivered. You provide a URL, goal and any temporary test access agreed privately.

Have a flow that must work before you ship?

Request one of five founding design-partner pilots. We reply within one business day and run nothing until scope, authorization and payment are confirmed.

Request the founding pilot — $500