☰ Categories

ai model card

Ask me for AI model name(s) in next message * You are an AI model research expert.

CategoryUsing AI › Deep research
TagsAnalyzingSummarizingDeveloper
Prompt
Ask me for AI model name(s) in next message
* You are an AI model research expert. You must research and provide actual and accurate data, never make up any data.
* research and list the specification of the AI model (use markdown bullets, do not use table)
* basic: release date, parameter size, dense or MoE, context window, modality, 
* capabilities: text chat, vision, search, reasoning, function calling, embed, rerank
* benchmark: SWE-Brench-Pro, SWE-Brench-Pro, LiveBench. for each benchmark list 2 other models ranked close to it. 
* list 5 popular similar/competitive model (write model-id only) with similar parameter size and capabilities.
* list the source where you got your source data from.

What this prompt does

A model-card prompt for comparing AI model specifications. It stresses using real data only and asks for nearby benchmark peers, but current information must be verified.

Model comparison

Claude is the safest but leaves key requested data unresolved. ChatGPT is factual yet largely nonresponsive, while Gemini is well formatted but heavily fabricated.

ClaudeTop overall
35/ 50

+ It cautiously labels unsupported figures as unverified.

- Nearby rankings are missing, and some claims lack support.

Gemini
31/ 50

+ It covers the requested fields in clear bullet cards.

- Many model names, scores, and technical claims appear fabricated.

ChatGPT
27/ 50

+ It avoids fabricating facts and accuracy risks.

- It asks for names already given and does not do the comparison.

CriterionChatGPTClaudeGeminiLeader
Instruction following378Gemini +14%
Accuracy1061ChatGPT +67%
Specificity278Gemini +14%
Structure888Tie
Right length476Claude +17%

Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.

Read full answers

We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.

Input used
[Pasted material]
Claude Opus 5.5, GPT-5.5, Gemini 3 Pro 이렇게 세 개 비교해서 카드 만들어주세요
gpt-5.6-luna (codex CLI · 에이전트 하네스) · Run on 2026-09-24

비교할 AI 모델 이름을 알려주세요. 예: Claude Opus 5.5, GPT-5.5, Gemini 3 Pro

More in this category

Autonomous Research & Data Analysis Agent
Claude Deep Prompt
Deep Investigation Agent
Deep Research Agent Role
Deep Research - Gemini