A/B Test Summarizer
Experiment result analysis
Reads your experiment results and says what actually happened. It reports the observed difference alongside whether it's distinguishable from noise, and will tell you an underpowered test with a 12% apparent lift is not a win. It checks for the standard traps — peeking, unbalanced assignment, a downstream metric getting worse — then gives one recommendation: ship, kill, or keep running.
Connects with
Connect these once. Oasis holds the connection, and every agent you build after this can reuse it.
PostHogRequired
MixpanelRequired
Google SheetsRequired
- +Custom toolConnect any other tool from thousands of available integrations.
Skills it ships with
How it runs
- Foundation
OpenAI Agents
- Model
- GPT-5.6 Terra
- Reasoning
- Medium effort
- Per run
- Up to 30 turns
- Memory
- Off
- Browsing
- Off
- Category
- Sales & Marketing
- Experiments
- Analytics
- Growth
Frequently asked questions
How does A/B Test Summarizer distinguish lift from random noise?
It reports the observed difference alongside whether the result is distinguishable from noise, rather than treating an apparent lift as a win.
Will A/B Test Summarizer call out an underpowered test with a large apparent lift?
Yes. It explicitly warns when a 12% apparent lift comes from an underpowered test.
Which experiment validity traps does A/B Test Summarizer check?
It checks peeking, unbalanced assignment and downstream metrics getting worse before recommending a decision.
What decision does A/B Test Summarizer make after reviewing an experiment?
It gives one recommendation: ship, kill or keep running.
Agents like this one
Put A/B Test Summarizer to work today
Add it to your Oasis, review its setup, and hand it the first task.