+ Conservatively limits claims to the supplied evidence.
- Supported-rung labels are somewhat vague and overcautious.
Climbs from spec to "so what" step by step, marking the rung where evidence becomes necessary.
| Category | Marketing › Ad copy |
|---|---|
| Tags | IdeationAnalyzingMarketer |
Build a benefit ladder from these features. For each feature, climb the rungs: 1. The feature as stated 2. What it does differently from the alternative 3. What that saves or enables for this customer 4. What that means for them beyond the product 5. Where it stops being defensible Output a table with all rungs, and mark: - The highest rung supported by the evidence I gave you. - The rung where a claim would need proof I do not have, with what proof would be needed. - Features that do not ladder at all — these belong in a spec list, not in copy. Rules: - *Do not climb past the evidence.* Rung 4 without proof is the line between marketing and a false claim, and it is where regulators and customers both look. - Ladders must be specific to the stated customer. The same feature ladders differently for different buyers. - Flag any rung that would be a regulated claim (health, financial, safety, efficacy) — those need review regardless of evidence.
Raw features do not sell; claims that climb too far become lies. This ascends one rung at a time and marks where a claim starts needing proof.
ChatGPT best respects the evidence boundary; Claude is the clearest and most actionable but makes some assumptions. Gemini is concrete yet invents multiple facts absent from the input.
+ Conservatively limits claims to the supplied evidence.
- Supported-rung labels are somewhat vague and overcautious.
+ Clearly marks support limits and missing proof.
- Assumes commute duration and some water-use benefits.
+ Offers concrete target scenarios and regulatory flags.
- Invents subway, competitor, commute, and preset data.
| Criterion | ChatGPT | Claude | Gemini | Leader |
|---|---|---|---|---|
| Instruction following | 9 | 10 | 7 | Claude +11% |
| Accuracy | 9 | 7 | 4 | ChatGPT +29% |
| Specificity | 9 | 10 | 8 | Claude +11% |
| Structure | 9 | 9 | 7 | Tie |
| Right length | 9 | 8 | 5 | ChatGPT +13% |
Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.
We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.
Features: 무선 이어폰 '사운드핏 프로' - 노이즈캔슬링 -32dB, 이어폰 8시간+케이스 포함 24시간 배터리, IPX4 방수, 전용 앱 이퀄라이저 지원 Target customer: 출퇴근길에 지하철을 이용하는 20~30대 직장인 Evidence I have: 자체 실험실 소음 차단 측정치(-32dB), 완충 후 실사용 배터리 테스트 결과, IPX4 방수 인증서. 소음 차단으로 인한 집중력·생산성 향상 관련 데이터는 보유하고 있지 않음
타깃을 “지하철로 출퇴근하는 2030대 직장인”으로 한정하면, 현재 근거만으로는 대부분 **제품 효용을 설명하는 13단계**까지만 안전합니다. 특히 집중력·생산성 향상은 현재 주장할 수 없습니다.
| 기능 | 1. 기능 자체 | 2. 대안과 다른 점 | 3. 고객에게 절약·가능하게 하는 것 | 4. 제품을 넘어 의미하는 것 | 5. 방어가 멈추는 지점 |
|---|---|---|---|---|---|
| 노이즈캔슬링 -32dB | 실험실 측정 기준 -32dB 소음 차단 | 일반 이어폰 또는 ANC가 없는 제품보다 지하철 소음을 더 줄이는 기능. 단, “더 줄인다”는 비교에는 대안 제품의 동일 조건 측정이 필요함 | 지하철 소음 속에서 음악·팟캐스트를 듣거나 업무에 필요한 소리를 듣기 쉬워질 가능성. 볼륨을 덜 높여도 된다는 주장은 추가 검증 필요 | “출퇴근 중 집중력과 생산성이 향상된다”는 의미는 현재 근거 없음 | 최고 입증 단계: 1단계. 2단계의 비교 주장에는 동일 환경·동일 측정법의 경쟁 제품 데이터가 필요. 3~4단계에는 실제 지하철 사용자 대상 청취성, 집중도, 업무수행 변화 데이터가 필요. 4단계는 효능성 주장으로 별도 검토 필요 |
| 배터리 8시간, 케이스 포함 24시간 | 이어폰 단독 최대 8시간, 케이스 포함 최대 24시간 사용 | 완충 후 실사용 테스트에서 해당 사용 시간이 확인됨. 대안보다 오래 간다는 표현에는 대안 배터리 데이터가 필요함 | 출퇴근 중 재충전 없이 사용할 가능성을 높이고, 충전 빈도를 줄일 수 있음. 다만 실제 출퇴근 시간과 사용 조건에 따라 달라짐 | “충전 걱정 없이 업무 루틴이 안정된다”는 표현은 사용자 조사나 실제 사용 데이터 없이는 과장될 수 있음 | 최고 입증 단계: 1단계~제한적 3단계. 테스트 조건·코덱·볼륨·통화 사용 여부를 명시해야 함. “하루 종일 충분하다”는 주장은 고객의 실제 사용시간 데이터가 필요 |
| IPX4 방수 | IPX4 방수 인증을 받음 | 방수 등급이 없거나 낮은 제품보다 생활 방수에 관한 명시적 보호 수준을 제공함. 비교하려면 대안의 등급 확인 필요 | 출퇴근 중 가벼운 비나 땀으로 인한 사용 불안을 줄일 수 있음. “비를 맞아도 안전하다” 또는 침수·고장 방지 보장은 아님 | “어떤 날씨에도 안심하고 사용할 수 있다”는 의미로 확장하면 과장 가능성이 큼 | 최고 입증 단계: 1단계~제한적 3단계. 인증서의 시험 범위와 IPX4의 한계를 함께 표시해야 함. 내구성·고장 방지·안전 보장은 별도 반복시험과 고장률 데이터가 필요 |
| 전용 앱 이퀄라이저 지원 | 전용 앱에서 이퀄라이저를 조절할 수 있음 | 고정된 음색만 제공하는 제품과 달리 사용자가 음색을 조정할 수 있음. 실제 앱 기능과 호환 기기 확인이 필요함 | 지하철처럼 소음이 큰 환경이나 음악 장르에 따라 원하는 청취 설정을 직접 만들 수 있음. “더 잘 들린다”는 주장은 청취 테스트가 필요함 | “내 취향과 업무 방식에 맞는 개인화된 경험을 제공한다” 정도는 메시지로 가능하지만, 만족도 향상은 조사 필요 | 최고 입증 단계: 1단계~제한적 3단계. 개인화가 실제 만족도·청취성·집중력 향상으로 이어진다는 주장은 사용자 테스트가 필요 |
| Build A/B test pairs that isolate one variable | |
| Write copy for different audience segments | |
| Define a brand tone and apply it | |
| Check ad copy for unsupported claims | |
| Write copy that contrasts with a competitor |