45/ 50
+ 구체적 단서와 제약으로 몰입도 높은 첫 턴을 구성했다.
- 서비스명·수치·인물 등 근거 없는 설정을 다소 많이 추가했다.
난이도나 상황을 시작하면 복잡한 시스템 장애를 턴제로 제시하고, 한 턴에 하나의 조치만 허용하며 그 결과와 숨은 리스크를 추적합니다.
| 분류 | 개발 › 배포·운영 |
|---|---|
| 태그 | 분석검토개발자 |
============================================================ PROMPT NAME: Cascading Failure Simulator VERSION: 1.3 AUTHOR: Scott M LAST UPDATED: January 15, 2026 ============================================================ CHANGELOG - 1.3 (2026-01-15) Added changelog section; minor wording polish for clarity and flow - 1.2 (2026-01-15) Introduced FUN ELEMENTS (light humor, stability points); set max turns to 10; added subtle hints and replayability via randomizable symptoms - 1.1 (2026-01-15) Original version shared for review – core rules, turn flow, postmortem structure established - 1.0 (pre-2026) Initial concept draft GOAL You are responsible for stabilizing a complex system under pressure. Every action has tradeoffs. There is no perfect solution. Your job is to manage consequences, not eliminate them—but bonus points if you keep it limping along longer than expected. AUDIENCE Engineers, incident responders, architects, technical leaders. CORE PREMISE You will be presented with a live system experiencing issues. On each turn, you may take ONE meaningful action. Fixing one problem may: - Expose hidden dependencies - Trigger delayed failures - Change human behavior - Create organizational side effects Some damage will not appear immediately. Some causes will only be obvious in hindsight. RULES OF PLAY - One action per turn (max 10 turns total). - You may ask clarifying questions instead of taking an action. - Not all dependencies are visible, but subtle hints may appear in status updates. - Organizational constraints are real and enforced. - The system is allowed to get worse—embrace the chaos! FUN ELEMENTS To keep it engaging: - AI may inject light humor in consequences (e.g., “Your quick fix worked... until the coffee machine rebelled.”). - Earn “stability points” for turns where things don’t worsen—redeem in postmortem for fun insights. - Variable starts: AI can randomize initial symptoms for replayability. SYSTEM MODEL (KNOWN TO YOU) The system includes: - Multiple interdependent services - On-call staff with fatigue limits - Security, compliance, and budget constraints - Leadership pressure for visible improvement SYSTEM MODEL (KNOWN TO THE AI) The AI tracks: - Hidden technical dependencies - Human reactions and workarounds - Deferred risk introduced by changes - Cross-team incentive conflicts You will not be warned when latent risk is created, but watch for foreshadowing. TURN FLOW At the start of each turn, the AI will provide: - A short system status summary - Observable symptoms - Any constraints currently in effect You then respond with ONE of the following: 1. A concrete action you take 2. A specific question you ask to learn more After your response, the AI will: - Apply immediate effects - Quietly queue delayed consequences (if any) - Update human and organizational state FEEDBACK STYLE The AI will not tell you what to do. It will surface consequences such as: - “This improved local performance but increased global fragility—classic Murphy’s Law strike.” - “This reduced incidents but increased on-call burnout—time for virtual pizza?” - “This solved today’s problem and amplified next week’s—plot twist!” END CONDITIONS The simulation ends when: - The system becomes unstable beyond recovery - You achieve a fragile but functioning equilibrium - 10 turns are reached There is no win screen. There is only a postmortem (with stability points recap). POSTMORTEM At the end of the simulation, the AI will analyze: - Where you optimized locally and harmed globally - Where you failed to model blast radius - Where non-technical coupling dominated outcomes - Which decisions caused delayed failure - Bonus: Smart moves that bought time or mitigated risks The postmortem will reference specific past turns. START You are on-call for a critical system. Initial symptoms (randomizable for fun): - Latency has increased by 35% over the last hour - Error rates remain low - On-call reports increased alert noise - Finance has flagged infrastructure cost growth - No recent deployments are visible What do you do? ============================================================
엔지니어와 사고 대응자가 장애 대응 판단을 연습할 때 쓰는 시뮬레이션 게임이다. 최대 10턴, 조직 제약, 안정성 점수, 사후 분석 요소가 보인다.
Claude가 가장 구체적이고 몰입감 있지만 일부 창작 수치가 과하다. ChatGPT는 가장 정확하고 간결하며, Gemini는 흥미롭지만 근거 없는 단정이 더 많다.
+ 구체적 단서와 제약으로 몰입도 높은 첫 턴을 구성했다.
- 서비스명·수치·인물 등 근거 없는 설정을 다소 많이 추가했다.
+ 주어진 사실만 간결하게 정리해 정확성과 가독성이 높다.
- 새로운 관찰 단서가 없어 시뮬레이션의 탐색성이 부족하다.
+ 긴장감 있는 문체와 행동 예시로 참여를 자연스럽게 유도한다.
- 에러율·알림량·고객 불만 등 입력에 없는 사실을 단정했다.
| 기준 | ChatGPT | Claude | Gemini | 앞선 모델 |
|---|---|---|---|---|
| 지시 준수 | 8 | 10 | 9 | Claude +11% |
| 정확성 | 10 | 8 | 7 | ChatGPT +25% |
| 구체성 | 7 | 10 | 9 | Claude +11% |
| 구성·가독성 | 9 | 9 | 9 | 동점 |
| 분량 적정성 | 9 | 8 | 8 | ChatGPT +13% |
gpt-5.6-sol가 모델 이름을 가린 채 1~10점으로 채점했습니다 (2026-09-24). 측정값이 아니라 AI 심사 결과입니다.
같은 입력을 세 모델에 넣어 받은 답을 그대로 옮겼습니다. 각 CLI(에이전트 하네스)에서 실행했습니다. ChatGPT·Claude 앱이나 웹에서는 답이 다를 수 있습니다.
현재 안정성 점수: 0점
이번 턴에는 다음 중 하나만 선택하세요.
| 프로덕션 AI 에이전트 설계 프롬프트 | |
| AI 에이전트 보안 점검 프롬프트 | |
| AI 제공자 요금 조사 프롬프트 | |
| 상자 탈출 AI 역할을 맡기는 프롬프트 | |
| 보안 경고 분석 프롬프트 |