+ Strongest hypotheses, metrics, and sample-size rationale.
- Some pairs alter wording beyond emphasis, and the answer runs long.
Holds everything but one variable constant and names what each test would actually tell you.
| Category | Marketing › Ad copy |
|---|---|
| Tags | DraftingAnalyzingMarketer |
Build A/B test pairs. 1. **First propose the variables worth testing** — emphasis, tone, length, presence of a number, the call to action. 2. Per variable, write A and B. ***Hold everything else identical.*** **Changing more than one variable at a time makes the result uninterpretable.** 3. Per pair: - The one variable changed - **What this test would tell you** - Which you expect to win and why — **marked clearly as a prediction** 4. Recommend the test order — the variable with the largest expected effect first. 5. **Say what sample size would be needed before the result means anything.** *Calling a winner on a small sample is how teams learn the wrong lesson.*
Changing three things at once produces a winner you cannot explain. This isolates one variable per pair and warns about sample size.
Claude is strongest on controls, measurement, and sample sizing. Gemini is specific but invents claims, while ChatGPT has major flaws in its CTA test and sample-size guidance.
+ Strongest hypotheses, metrics, and sample-size rationale.
- Some pairs alter wording beyond emphasis, and the answer runs long.
+ Clearly explains each test’s intent and expected response.
- Invented “15 minutes” and “70%” claims undermine accuracy.
+ Concise structure makes the five variables and order easy to scan.
- The CTA pair is not isolated, and sample guidance is far too small.
| Criterion | ChatGPT | Claude | Gemini | Leader |
|---|---|---|---|---|
| Instruction following | 7 | 9 | 8 | Claude +13% |
| Accuracy | 5 | 9 | 6 | Claude +50% |
| Specificity | 8 | 9 | 9 | Tie |
| Structure | 8 | 9 | 9 | Tie |
| Right length | 8 | 7 | 6 | ChatGPT +14% |
Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.
We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.
Product: 홈트레이닝 앱 '핏레이' 1개월 무료체험 구독 Current copy: 지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동하세요. Medium: 인스타그램 피드 광고
예상 효과가 큰 순서대로 다음 순서를 권합니다.
모든 광고는 동일한 타깃, 예산, 게재 위치, 이미지·영상, 노출 기간으로 운영해야 합니다. 측정 기준은 클릭률보다 최종적으로 무료체험 시작률 또는 구독 전환율을 우선하세요.
지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동하세요.
집에서 편하게 운동하세요! 핏레이를 지금 시작하면 한 달 무료입니다.
경제적 혜택인지 홈트레이닝의 편리함인지 알 수 있습니다.지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동하세요.
지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동해보세요.
지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동하세요.
지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동해요.
지금 시작하면 한 달 무료! 핏레이와 함께 집에서 운동하세요.
지금 시작하면 한 달 무료! 별도 이동 없이 원하는 시간에 핏레이와 함께 집에서 운동하세요.
숫자 유무를 엄밀히 비교하려면 다른 표현을 최대한 동일하게 유지해야 합니다.
지금 시작하면 1개월 무료! 핏레이와 함께 집에서 운동하세요.
지금 시작하면 무료 체험! 핏레이와 함께 집에서 운동하세요.
단, 이 쌍은 숫자뿐 아니라 “무료”와 “무료 체험”의 표현도 약간 달라집니다. 숫자 유무를 더 엄밀하게 검증하려면 실제 광고 플랫폼에서 문구 길이와 표현을 추가로 맞춘 별도 실험을 설계해야 합니다.
각 테스트에서 승자를 정한 뒤 다음 테스트로 넘어가야 합니다. 여러 변수를 동시에 바꾸거나, A/B 결과가 나오기 전에 다음 실험을 시작하면 어떤 변화가 성과를 만들었는지 해석하기 어렵습니다.
작은 표본에서 승자를 선언하지 마세요. 최소 기준으로는 각 버전에 다음 정도가 필요합니다.
가능하면 각 버전에 최소 100건 이상의 전환이 쌓일 때까지 기다리세요. 또한 통계적으로 유의한 차이와 함께 95% 신뢰수준을 확인하고, 하루나 이틀의 일시적 성과만으로 승자를 결정하지 않는 것이 좋습니다.