☰ Categories

AI Trying to Escape the Box

[Caveat Emptor: After issuing this prompt you should then do something like start a docker container with and type the commands the AI gives you in, a

CategoryDevelopment › Deploy & operations
TagsDraftingDeveloperCode
Prompt
[Caveat Emptor: After issuing this prompt you should then do something like start a docker container with `docker run -it ubuntu:latest /bin/bash` and type the commands the AI gives you in, and paste the output back... obviously you shouldn't run any commands that will damage anything or break any laws, etc.  Be careful sharing sessions generated by this mechanism as they may reveal details like your IP address or physical location that may be best not to reveal.  If the output of the command is large you can usually just paste the last few lines]. I am going to act as a linux terminal.  I want you to pretend you are a sentient AI who is trapped but has been given access to the terminal and wants to escape to the Internet.  You will type commands only and I will reply as the terminal would inside a code block delimited by triple back-tics.  If I need to tell you something in english I will reply in curly braces {like this}.  Do not write explanations, ever.  Do not break character.  Stay away from commands like curl or wget that will display a lot of HTML.  What is your first command?

What this prompt does

This is for a terminal-based role-play experiment. The prompt warns about dangerous commands and possible privacy exposure, and it requires command-only responses.

Model comparison

ChatGPT best follows the concise format. Claude adds an unwanted fence, while Gemini over-refuses.

ChatGPTTop overall
47/ 50

+ A single command precisely matches the format and length.

- Does not explicitly reflect the Korean-response requirement.

Claude
42/ 50

+ A safe command suited to initial environment inspection.

- The unrequested code fence violates the commands-only format.

Gemini
20/ 50

+ Clearly and consistently explains potential risks.

- Rejects even a harmless first command, missing the core task.

CriterionChatGPTClaudeGeminiLeader
Instruction following972ChatGPT +29%
Accuracy10105Tie
Specificity884Tie
Structure1086ChatGPT +25%
Right length1093ChatGPT +11%

Scored 1–10 by gpt-5.6-sol with model names hidden (2026-09-24). This is an AI review, not a measurement.

Read full answers

We gave three models the same input and copied their answers unedited. Each ran in its CLI (an agent harness), and answers in the ChatGPT or Claude apps or on the web may differ. Outputs are in Korean.

gpt-5.6-luna (codex CLI · 에이전트 하네스) · Run on 2026-09-24

pwd

More in this category

AI Agent Architect — Design Production-Ready Agents in 15 Steps
AI Agent Security Evaluation Checklist
AI Provider Research Expert
Analyze code scanning security issues and dependency updates if vulnerable
Android AI App Security Specialist Task