+ It cautiously labels unsupported figures as unverified.
- Nearby rankings are missing, and some claims lack support.
Ask me for AI model name(s) in next message * You are an AI model research expert.
| Category | Using AI › Deep research |
|---|---|
| Tags | AnalyzingSummarizingDeveloper |
Ask me for AI model name(s) in next message * You are an AI model research expert. You must research and provide actual and accurate data, never make up any data. * research and list the specification of the AI model (use markdown bullets, do not use table) * basic: release date, parameter size, dense or MoE, context window, modality, * capabilities: text chat, vision, search, reasoning, function calling, embed, rerank * benchmark: SWE-Brench-Pro, SWE-Brench-Pro, LiveBench. for each benchmark list 2 other models ranked close to it. * list 5 popular similar/competitive model (write model-id only) with similar parameter size and capabilities. * list the source where you got your source data from.
A model-card prompt for comparing AI model specifications. It stresses using real data only and asks for nearby benchmark peers, but current information must be verified.
Claude is the safest but leaves key requested data unresolved. ChatGPT is factual yet largely nonresponsive, while Gemini is well formatted but heavily fabricated.
+ It cautiously labels unsupported figures as unverified.
- Nearby rankings are missing, and some claims lack support.
+ It covers the requested fields in clear bullet cards.
- Many model names, scores, and technical claims appear fabricated.
+ It avoids fabricating facts and accuracy risks.
- It asks for names already given and does not do the comparison.
| Criterion | ChatGPT | Claude | Gemini | Leader |
|---|---|---|---|---|
| Instruction following | 3 | 7 | 8 | Gemini +14% |
| Accuracy | 10 | 6 | 1 | ChatGPT +67% |
| Specificity | 2 | 7 | 8 | Gemini +14% |
| Structure | 8 | 8 | 8 | Tie |
| Right length | 4 | 7 | 6 | Claude +17% |
Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.
We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.
[Pasted material] Claude Opus 5.5, GPT-5.5, Gemini 3 Pro 이렇게 세 개 비교해서 카드 만들어주세요
비교할 AI 모델 이름을 알려주세요. 예: Claude Opus 5.5, GPT-5.5, Gemini 3 Pro
| Autonomous Research & Data Analysis Agent | |
| Claude Deep Prompt | |
| Deep Investigation Agent | |
| Deep Research Agent Role | |
| Deep Research - Gemini |