imagegen skill (openai/skills)
- Install
- SKILL.md (verbatim)
- Top-level modes and rules
- When to use
- When not to use
- Decision tree
- Workflow
- Prompt augmentation
- Specificity policy
- Use-case taxonomy (exact slugs)
- Shared prompt schema
- Examples
- Generation example (hero image)
- Edit example (invariants)
- Prompting best practices
- Guidance by asset type
- Fallback CLI mode only
- Temp and output conventions
- Dependencies
- Environment
- Script-mode notes
- Reference map
- Other files in this skill
- references/cli.md (verbatim)
- What this CLI does
- Quick start (works from any repo)
- Quick start
- Guardrails
- Defaults
- Quality, input fidelity, and masks (CLI fallback only)
- Output handling
- Common recipes
- CLI notes
- See also
- references/codex-network.md (verbatim)
- Why am I asked to approve image generation calls?
- Important note about approvals vs network
- How do I reduce repeated approval prompts?
- Safety note
- references/image-api.md (verbatim)
- Scope
- Endpoints
- Core parameters for GPT Image models
- Edit-specific parameters
- Output
- Limits and notes
- Important boundary
- references/prompting.md (verbatim)
- Contents
- Structure
- Specificity policy
- Allowed and disallowed augmentation
- Composition and layout
- Constraints and invariants
- Text in images
- Input images and references
- Iterate deliberately
- Fallback-only execution controls
- Use-case tips
- Where to find copy/paste recipes
- references/sample-prompts.md (verbatim)
- Generate
- photorealistic-natural
- product-mockup
- ui-mockup
- infographic-diagram
- logo-brand
- illustration-story
- stylized-concept
- historical-scene
- Asset type templates (taxonomy-aligned)
- Website assets template
- Website assets example: minimal hero background
- Website assets example: feature section illustration
- Website assets example: blog header image
- Game assets template
- Game assets example: environment concept art
- Game assets example: character concept
- Game assets example: UI icon
- Game assets example: tileable texture
- Wireframe template
- Wireframe example: homepage (desktop)
- Wireframe example: pricing page
- Wireframe example: mobile onboarding flow
- Logo template
- Logo example: abstract symbol mark
- Logo example: monogram mark
- Logo example: wordmark
- Edit
- text-localization
- identity-preserve
- precise-object-edit
- lighting-weather
- background-extraction
- style-transfer
- compositing
- sketch-to-render
What it does. Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas. Part of openai/skills (Skills Catalog for Codex) (openai/skills).
| Upstream | openai/skills |
| Skill file | skills/.system/imagegen/SKILL.md |
| License | Apache-2.0 (skill folder LICENSE.txt) |
| Author | OpenAI |
| Fetched | 2026-09-10 |
Install
- Codex:
$skill-installerinstalls from this catalog (this one ships with Codex by default); other agents:npx skills add openai/skills --skill imagegen. - Raw file:
curl -sL https://raw.githubusercontent.com/openai/skills/HEAD/skills/.system/imagegen/SKILL.md
SKILL.md (verbatim)
name: "imagegen"
description: "Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output should be a bitmap asset rather than repo-native code or vector. Do not use when the task is better handled by editing existing SVG/vector/code-native assets, extending an established icon or logo system, or building the visual directly in HTML/CSS/canvas."
Image Generation Skill
Generates or edits images for the current project (for example website assets, game assets, UI mockups, product mockups, wireframes, logo design, photorealistic images, or infographics).
Top-level modes and rules
This skill has exactly two top-level modes:
- Default built-in tool mode (preferred): built-in
image_gentool for normal image generation and editing. Does not requireOPENAI_API_KEY. - Fallback CLI mode (explicit-only):
scripts/image_gen.pyCLI. Use only when the user explicitly asks for the CLI path. RequiresOPENAI_API_KEY.
Within the explicit CLI fallback only, the CLI exposes three subcommands:
generateeditgenerate-batch
Rules:
- Use the built-in
image_gentool by default for all normal image generation and editing requests. - Never switch to CLI fallback automatically.
- If the built-in tool fails or is unavailable, tell the user the CLI fallback exists and that it requires
OPENAI_API_KEY. Proceed only if the user explicitly asks for that fallback. - If the user explicitly asks for CLI mode, use the bundled
scripts/image_gen.pyworkflow. Do not create one-off SDK runners. - Never modify
scripts/image_gen.py. If something is missing, ask the user before doing anything else.
Built-in save-path policy:
- In built-in tool mode, Codex saves generated images under
$CODEX_HOME/*by default. - Do not describe or rely on OS temp as the default built-in destination.
- Do not describe or rely on a destination-path argument (if any) on the built-in
image_gentool. If a specific location is needed, generate first and then move or copy the selected output from$CODEX_HOME/generated_images/.... - Save-path precedence in built-in mode:
- If the user names a destination, move or copy the selected output there.
- If the image is meant for the current project, move or copy the final selected image into the workspace before finishing.
- If the image is only for preview or brainstorming, render it inline; the underlying file can remain at the default
$CODEX_HOME/*path.
- Never leave a project-referenced asset only at the default
$CODEX_HOME/*path. - Do not overwrite an existing asset unless the user explicitly asked for replacement; otherwise create a sibling versioned filename such as
hero-v2.pngoritem-icon-edited.png.
Shared prompt guidance for both modes lives in references/prompting.md and references/sample-prompts.md.
Fallback-only docs/resources for CLI mode:
references/cli.mdreferences/image-api.mdreferences/codex-network.mdscripts/image_gen.py
When to use
- Generate a new image (concept art, product shot, cover, website hero)
- Generate a new image using one or more reference images for style, composition, or mood
- Edit an existing image (inpainting, lighting or weather transformations, background replacement, object removal, compositing, transparent background)
- Produce many assets or variants for one task
When not to use
- Extending or matching an existing SVG/vector icon set, logo system, or illustration library inside the repo
- Creating simple shapes, diagrams, wireframes, or icons that are better produced directly in SVG, HTML/CSS, or canvas
- Making a small project-local asset edit when the source file already exists in an editable native format
- Any task where the user clearly wants deterministic code-native output instead of a generated bitmap
Decision tree
Think about two separate questions:
- Intent: is this a new image or an edit of an existing image?
- Execution strategy: is this one asset or many assets/variants?
Intent:
- If the user wants to modify an existing image while preserving parts of it, treat the request as edit.
- If the user provides images only as references for style, composition, mood, or subject guidance, treat the request as generate.
- If the user provides no images, treat the request as generate.
Built-in edit semantics:
- Built-in edit mode is for images already visible in the conversation context, such as attached images or images generated earlier in the thread.
- If the user wants to edit a local image file with the built-in tool, first load it with built-in
view_imagetool so the image is visible in the conversation context, then proceed with the built-in edit flow. - Do not promise arbitrary filesystem-path editing through the built-in tool.
- If a local file still needs direct file-path control, masks, or other explicit CLI-only parameters, use the explicit CLI fallback only when the user asks for it.
- For edits, preserve invariants aggressively and save non-destructively by default.
Execution strategy:
- In the built-in default path, produce many assets or variants by issuing one
image_gencall per requested asset or variant. - In the explicit CLI fallback path, use the CLI
generate-batchsubcommand only when the user explicitly chose CLI mode and needs many prompts/assets.
Assume the user wants a new image unless they clearly ask to change an existing one.
Workflow
- Decide the top-level mode: built-in by default, fallback CLI only if explicitly requested.
- Decide the intent:
generateoredit. - Decide whether the output is preview-only or meant to be consumed by the current project.
- Decide the execution strategy: single asset vs repeated built-in calls vs CLI
generate-batch. - Collect inputs up front: prompt(s), exact text (verbatim), constraints/avoid list, and any input images.
- For every input image, label its role explicitly:
- reference image
- edit target
- supporting insert/style/compositing input
- If the edit target is only on the local filesystem and you are staying on the built-in path, inspect it with
view_imagefirst so the image is available in conversation context. - If the user asked for a photo, illustration, sprite, product image, banner, or other explicitly raster-style asset, use
image_genrather than substituting SVG/HTML/CSS placeholders. If the request is for an icon, logo, or UI graphic that should match existing repo-native SVG/vector/code assets, prefer editing those directly instead. - Augment the prompt based on specificity:
- If the user's prompt is already specific and detailed, normalize it into a clear spec without adding creative requirements.
- If the user's prompt is generic, add tasteful augmentation only when it materially improves output quality.
- Use the built-in
image_gentool by default. - If the user explicitly chooses the CLI fallback, then and only then use the fallback-only docs for quality,
input_fidelity, masks, output format, output paths, and network setup. - Inspect outputs and validate: subject, style, composition, text accuracy, and invariants/avoid items.
- Iterate with a single targeted change, then re-check.
- For preview-only work, render the image inline; the underlying file may remain at the default
$CODEX_HOME/generated_images/...path. - For project-bound work, move or copy the selected artifact into the workspace and update any consuming code or references. Never leave a project-referenced asset only at the default
$CODEX_HOME/generated_images/...path. - For batches, persist only the selected finals in the workspace unless the user explicitly asked to keep discarded variants.
- Always report the final saved path for any workspace-bound asset, plus the final prompt and whether the built-in tool or fallback CLI mode was used.
Prompt augmentation
Reformat user prompts into a structured, production-oriented spec. Make the user's goal clearer and more actionable, but do not blindly add detail.
Treat this as prompt-shaping guidance, not a closed schema. Use only the lines that help, and add a short extra labeled line when it materially improves clarity.
Specificity policy
Use the user's prompt specificity to decide how much augmentation is appropriate:
- If the prompt is already specific and detailed, preserve that specificity and only normalize/structure it.
- If the prompt is generic, you may add tasteful augmentation when it will materially improve the result.
Allowed augmentations:
- composition or framing hints
- polish level or intended-use hints
- practical layout guidance
- reasonable scene concreteness that supports the stated request
Not allowed augmentations:
- extra characters or objects that are not implied by the request
- brand names, slogans, palettes, or narrative beats that are not implied
- arbitrary side-specific placement unless the surrounding layout supports it
Use-case taxonomy (exact slugs)
Classify each request into one of these buckets and keep the slug consistent across prompts and references.
Generate:
- photorealistic-natural — candid/editorial lifestyle scenes with real texture and natural lighting.
- product-mockup — product/packaging shots, catalog imagery, merch concepts.
- ui-mockup — app/web interface mockups and wireframes; specify the desired fidelity.
- infographic-diagram — diagrams/infographics with structured layout and text.
- logo-brand — logo/mark exploration, vector-friendly.
- illustration-story — comics, children’s book art, narrative scenes.
- stylized-concept — style-driven concept art, 3D/stylized renders.
- historical-scene — period-accurate/world-knowledge scenes.
Edit:
- text-localization — translate/replace in-image text, preserve layout.
- identity-preserve — try-on, person-in-scene; lock face/body/pose.
- precise-object-edit — remove/replace a specific element (including interior swaps).
- lighting-weather — time-of-day/season/atmosphere changes only.
- background-extraction — transparent background / clean cutout.
- style-transfer — apply reference style while changing subject/scene.
- compositing — multi-image insert/merge with matched lighting/perspective.
- sketch-to-render — drawing/line art to photoreal render.
Shared prompt schema
Use the following labeled spec as shared prompt scaffolding for both top-level modes:
Use case: <taxonomy slug>
Asset type: <where the asset will be used>
Primary request: <user's main prompt>
Input images: <Image 1: role; Image 2: role> (optional)
Scene/backdrop: <environment>
Subject: <main subject>
Style/medium: <photo/illustration/3D/etc>
Composition/framing: <wide/close/top-down; placement>
Lighting/mood: <lighting + mood>
Color palette: <palette notes>
Materials/textures: <surface details>
Text (verbatim): "<exact text>"
Constraints: <must keep/must avoid>
Avoid: <negative constraints>
Notes:
Asset typeandInput imagesare prompt scaffolding, not dedicated CLI flags.Scene/backdroprefers to the visual setting. It is not the same as the fallback CLIbackgroundparameter, which controls output transparency behavior.- Fallback-only execution notes such as
Quality:,Input fidelity:, masks, output format, and output paths belong in the explicit CLI path only. Do not treat them as built-inimage_gentool arguments.
Augmentation rules:
- Keep it short.
- Add only the details needed to improve the prompt materially.
- For edits, explicitly list invariants (
change only X; keep Y unchanged). - If any critical detail is missing and blocks success, ask a question; otherwise proceed.
Examples
Generation example (hero image)
Use case: product-mockup
Asset type: landing page hero
Primary request: a minimal hero image of a ceramic coffee mug
Style/medium: clean product photography
Composition/framing: wide composition with usable negative space for page copy if needed
Lighting/mood: soft studio lighting
Constraints: no logos, no text, no watermark
Edit example (invariants)
Use case: precise-object-edit
Asset type: product photo background replacement
Primary request: replace only the background with a warm sunset gradient
Constraints: change only the background; keep the product and its edges unchanged; no text; no watermark
Prompting best practices
- Structure prompt as scene/backdrop -> subject -> details -> constraints.
- Include intended use (ad, UI mock, infographic) to set the mode and polish level.
- Use camera/composition language for photorealism.
- Only use SVG/vector stand-ins when the user explicitly asked for vector output or a non-image placeholder.
- Quote exact text and specify typography + placement.
- For tricky words, spell them letter-by-letter and require verbatim rendering.
- For multi-image inputs, reference images by index and describe how they should be used.
- For edits, repeat invariants every iteration to reduce drift.
- Iterate with single-change follow-ups.
- If the prompt is generic, add only the extra detail that will materially help.
- If the prompt is already detailed, normalize it instead of expanding it.
- For explicit CLI fallback only, see
references/cli.mdandreferences/image-api.mdforquality,input_fidelity, masks, output format, and output-path guidance.
More principles shared by both modes: references/prompting.md.
Copy/paste specs shared by both modes: references/sample-prompts.md.
Guidance by asset type
Asset-type templates (website assets, game assets, wireframes, logo) are consolidated in references/sample-prompts.md.
Fallback CLI mode only
Temp and output conventions
These conventions apply only to the explicit CLI fallback. They do not describe built-in image_gen output behavior.
- Use
tmp/imagegen/for intermediate files (for example JSONL batches); delete them when done. - Write final artifacts under
output/imagegen/. - Use
--outor--out-dirto control output paths; keep filenames stable and descriptive.
Dependencies
Prefer uv for dependency management in this repo.
Required Python package:
uv pip install openai
Optional for downscaling only:
uv pip install pillow
Portability note:
- If you are using the installed skill outside this repo, install dependencies into that environment with its package manager.
- In uv-managed environments,
uv pip install ...remains the preferred path.
Environment
OPENAI_API_KEYmust be set for live API calls.- Do not ask the user for
OPENAI_API_KEYwhen using the built-inimage_gentool. - Never ask the user to paste the full key in chat. Ask them to set it locally and confirm when ready.
If the key is missing, give the user these steps:
- Create an API key in the OpenAI platform UI: https://platform.openai.com/api-keys
- Set
OPENAI_API_KEYas an environment variable in their system. - Offer to guide them through setting the environment variable for their OS/shell if needed.
If installation is not possible in this environment, tell the user which dependency is missing and how to install it into their active environment.
Script-mode notes
- CLI commands + examples:
references/cli.md - API parameter quick reference:
references/image-api.md - Network approvals / sandbox settings for CLI mode:
references/codex-network.md
Reference map
references/prompting.md: shared prompting principles for both modes.references/sample-prompts.md: shared copy/paste prompt recipes for both modes.references/cli.md: fallback-only CLI usage viascripts/image_gen.py.references/image-api.md: fallback-only API/CLI parameter reference.references/codex-network.md: fallback-only network/sandbox troubleshooting for CLI mode.scripts/image_gen.py: fallback-only CLI implementation. Do not load or use it unless the user explicitly chooses CLI mode.
Other files in this skill
- LICENSE.txt
- agents/openai.yaml
- assets/imagegen-small.svg
- assets/imagegen.png
- references/cli.md
- references/codex-network.md
- references/image-api.md
- references/prompting.md
- references/sample-prompts.md
- scripts/image_gen.py
references/cli.md (verbatim)
CLI reference (scripts/image_gen.py)
This file is for the fallback CLI mode only. Read it only after the user explicitly asks to use scripts/image_gen.py instead of the built-in image_gen tool.
generate-batch is a CLI subcommand in this fallback path. It is not a top-level mode of the skill.
What this CLI does
generate: generate a new image from a promptedit: edit one or more existing imagesgenerate-batch: run many generation jobs from a JSONL file
Real API calls require network access + OPENAI_API_KEY. --dry-run does not.
Quick start (works from any repo)
Set a stable path to the skill CLI (default CODEX_HOME is ~/.codex):
export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export IMAGE_GEN="$CODEX_HOME/skills/.system/imagegen/scripts/image_gen.py"
Install dependencies into that environment with its package manager. In uv-managed environments, uv pip install ... remains the preferred path.
Quick start
Dry-run (no API call; no network required; does not require the openai package):
python "$IMAGE_GEN" generate \
--prompt "Test" \
--out output/imagegen/test.png \
--dry-run
Notes:
- One-off dry-runs print the API payload and the computed output path(s).
- Repo-local finals should live under
output/imagegen/.
Generate (requires OPENAI_API_KEY + network):
python "$IMAGE_GEN" generate \
--prompt "A cozy alpine cabin at dawn" \
--size 1024x1024 \
--out output/imagegen/alpine-cabin.png
Edit:
python "$IMAGE_GEN" edit \
--image input.png \
--prompt "Replace only the background with a warm sunset" \
--out output/imagegen/sunset-edit.png
Guardrails
- Use the bundled CLI directly (
python "$IMAGE_GEN" ...) after activating the correct environment. - Do not create one-off runners (for example
gen_images.py) unless the user explicitly asks for a custom wrapper. - Never modify
scripts/image_gen.py. If something is missing, ask the user before doing anything else.
Defaults
- Model:
gpt-image-1.5 - Supported model family for this CLI: GPT Image models (
gpt-image-*) - Size:
1024x1024 - Quality:
auto - Output format:
png - Default one-off output path:
output/imagegen/output.png - Background: unspecified unless
--backgroundis set
Quality, input fidelity, and masks (CLI fallback only)
These are explicit CLI controls. They are not built-in image_gen tool arguments.
--qualityworks forgenerate,edit, andgenerate-batch:low|medium|high|auto--input-fidelityis edit-only and validated aslow|high--maskis edit-only
Example:
python "$IMAGE_GEN" edit \
--image input.png \
--prompt "Change only the background" \
--quality high \
--input-fidelity high \
--out output/imagegen/background-edit.png
Mask notes:
- For multi-image edits, pass repeated
--imageflags. Their order is meaningful, so describe each image by index and role in the prompt. - The CLI accepts a single
--mask. - Use a PNG mask when possible; the script treats mask handling as best-effort and does not perform full preflight validation beyond file checks/warnings.
- In the edit prompt, repeat invariants (
change only the background; keep the subject unchanged) to reduce drift.
Output handling
- Use
tmp/imagegen/for temporary JSONL inputs or scratch files. - Use
output/imagegen/for final outputs. - Reruns fail if a target file already exists unless you pass
--force. --out-dirchanges one-off naming toimage_1.<ext>,image_2.<ext>, and so on.- Downscaled copies use the default suffix
-webunless you override it.
Common recipes
Generate with augmentation fields:
python "$IMAGE_GEN" generate \
--prompt "A minimal hero image of a ceramic coffee mug" \
--use-case "product-mockup" \
--style "clean product photography" \
--composition "wide product shot with usable negative space for page copy" \
--constraints "no logos, no text" \
--out output/imagegen/mug-hero.png
Generate + also write a downscaled copy for fast web loading:
python "$IMAGE_GEN" generate \
--prompt "A cozy alpine cabin at dawn" \
--size 1024x1024 \
--downscale-max-dim 1024 \
--out output/imagegen/alpine-cabin.png
Generate multiple prompts concurrently (async batch):
mkdir -p tmp/imagegen output/imagegen/batch
cat > tmp/imagegen/prompts.jsonl << 'EOF'
{"prompt":"Cavernous hangar interior with a compact shuttle parked near the center","use_case":"stylized-concept","composition":"wide-angle, low-angle","lighting":"volumetric light rays through drifting fog","constraints":"no logos or trademarks; no watermark","size":"1536x1024"}
{"prompt":"Gray wolf in profile in a snowy forest","use_case":"photorealistic-natural","composition":"eye-level","constraints":"no logos or trademarks; no watermark","size":"1024x1024"}
EOF
python "$IMAGE_GEN" generate-batch \
--input tmp/imagegen/prompts.jsonl \
--out-dir output/imagegen/batch \
--concurrency 5
rm -f tmp/imagegen/prompts.jsonl
Notes:
generate-batchrequires--out-dir.- generate-batch requires --out-dir.
- Use
--concurrencyto control parallelism (default5). - Per-job overrides are supported in JSONL (for example
size,quality,background,output_format,output_compression,moderation,n,model,out, and prompt-augmentation fields). --ngenerates multiple variants for a single prompt;generate-batchis for many different prompts.- In batch mode, per-job
outis treated as a filename under--out-dir.
CLI notes
- Supported sizes:
1024x1024,1536x1024,1024x1536, orauto. - Transparent backgrounds require
output_formatto bepngorwebp. --prompt-file,--output-compression,--moderation,--max-attempts,--fail-fast,--force, and--no-augmentare supported.- This CLI is intended for GPT Image models. Do not assume older non-GPT image-model behavior applies here.
See also
- API parameter quick reference for fallback CLI mode:
references/image-api.md - Prompt examples shared across both top-level modes:
references/sample-prompts.md - Network/sandbox notes for fallback CLI mode:
references/codex-network.md
references/codex-network.md (verbatim)
Codex network approvals / sandbox notes
This file is for the fallback CLI mode only. Read it only after the user explicitly asks to use scripts/image_gen.py.
This guidance is intentionally isolated from SKILL.md because it can vary by environment and may become stale. Prefer the defaults in your environment when in doubt.
Why am I asked to approve image generation calls?
The fallback CLI uses the OpenAI Image API, so it needs outbound network access. In many Codex setups, network access is disabled by default and/or the approval policy requires confirmation before networked commands run.
Important note about approvals vs network
--ask-for-approval neversuppresses approval prompts.- It does not by itself enable network access.
- In
workspace-write, network access still depends on your Codex configuration (for example[sandbox_workspace_write] network_access = true).
How do I reduce repeated approval prompts?
If you trust the repo and want fewer prompts, use a configuration or profile that both:
- enables network for the sandbox mode you plan to use
- sets an approval policy that matches your risk tolerance
Example ~/.codex/config.toml pattern:
approval_policy = "on-request"
sandbox_mode = "workspace-write"
[sandbox_workspace_write]
network_access = true
If you want quieter automation after network is enabled, you can choose a stricter approval policy, but do that intentionally and with care.
Safety note
Enabling network and reducing approvals lowers friction, but increases risk if you run untrusted code or work in an untrusted repository.
references/image-api.md (verbatim)
Image API quick reference
This file is for the fallback CLI mode only. Use it only after the user explicitly asks to use scripts/image_gen.py instead of the built-in image_gen tool.
These parameters describe the Image API and bundled CLI fallback surface. Do not assume they are normal arguments on the built-in image_gen tool.
Scope
- This fallback CLI is intended for GPT Image models (
gpt-image-1.5,gpt-image-1, andgpt-image-1-mini). - The built-in
image_gentool and the fallback CLI do not expose the same controls.
Endpoints
- Generate:
POST /v1/images/generations(client.images.generate(...)) - Edit:
POST /v1/images/edits(client.images.edit(...))
Core parameters for GPT Image models
prompt: text promptmodel: image modeln: number of images (1-10)size:1024x1024,1536x1024,1024x1536, orautoquality:low,medium,high, orautobackground: output transparency behavior (transparent,opaque, orauto) for generated output; this is not the same thing as the prompt's visual scene/backdropoutput_format:png(default),jpeg,webpoutput_compression: 0-100 (jpeg/webp only)moderation:auto(default) orlow
Edit-specific parameters
image: one or more input images. For GPT Image models, you can provide up to 16 images.mask: optional mask imageinput_fidelity:low(default) orhigh
Model-specific note for input_fidelity:
gpt-image-1andgpt-image-1-minipreserve all input images, but the first image gets richer textures and finer details.gpt-image-1.5preserves the first 5 input images with higher fidelity.
Output
data[]list withb64_jsonper image- The bundled
scripts/image_gen.pyCLI decodesb64_jsonand writes output files for you.
Limits and notes
- Input images and masks must be under 50MB.
- Use the edits endpoint when the user requests changes to an existing image.
- Masking is prompt-guided; exact shapes are not guaranteed.
- Large sizes and high quality increase latency and cost.
- High
input_fidelitycan materially increase input token usage. - If a request fails because a specific option is unsupported by the selected GPT Image model, retry manually without that option.
Important boundary
quality,input_fidelity, explicit masks,background,output_format, and related parameters are fallback-only execution controls.- Do not assume they are built-in
image_gentool arguments.
references/prompting.md (verbatim)
Prompting best practices
These prompting principles are shared by both top-level modes of the skill:
- built-in
image_gentool (default) - explicit
scripts/image_gen.pyCLI fallback
This file is about prompt structure, specificity, and iteration. Fallback-only execution controls such as quality, input_fidelity, masks, output format, and output paths live in the fallback docs.
Contents
- Structure
- Specificity policy
- Allowed and disallowed augmentation
- Composition and layout
- Constraints and invariants
- Text in images
- Input images and references
- Iterate deliberately
- Fallback-only execution controls
- Use-case tips
- Where to find copy/paste recipes
Structure
- Use a consistent order: scene/backdrop -> subject -> key details -> constraints -> output intent.
- Include intended use (ad, UI mock, infographic) to set the level of polish.
- For complex requests, use short labeled lines instead of one long paragraph.
Specificity policy
- If the user prompt is already specific and detailed, normalize it into a clean spec without adding creative requirements.
- If the prompt is generic, you may add tasteful detail when it materially improves the output.
- Treat examples in
sample-prompts.mdas fully-authored recipes, not as the default amount of augmentation to add to every request.
Allowed and disallowed augmentation
Allowed augmentation for generic prompts:
- composition and framing cues
- intended-use or polish-level hints
- practical layout guidance
- reasonable scene concreteness that supports the request
Do not add:
- extra characters, props, or objects that are not implied
- brand palettes, slogans, or story beats that are not implied
- arbitrary side-specific placement unless the surrounding layout supports it
Composition and layout
- Specify framing and viewpoint (close-up, wide, top-down) and placement only when it materially helps.
- Call out negative space if the asset clearly needs room for UI or copy.
- Avoid making left/right layout decisions unless the user or surrounding layout supports them.
Constraints and invariants
- State what must not change (
keep background unchanged). - For edits, say
change only X; keep Y unchangedand repeat invariants on every iteration to reduce drift.
Text in images
- Put literal text in quotes or ALL CAPS and specify typography (font style, size, color, placement).
- Spell uncommon words letter-by-letter if accuracy matters.
- For in-image copy, require verbatim rendering and no extra characters.
Input images and references
- Do not assume that every provided image is an edit target.
- Label each image by index and role (
Image 1: edit target,Image 2: style reference). - If the user provides images for style, composition, or mood guidance and does not ask to modify them, treat the request as generation with references.
- If the user asks to preserve an existing image while changing specific parts, treat the request as an edit.
- For compositing, describe how the images interact (
place the subject from Image 2 into Image 1).
Iterate deliberately
- Start with a clean base prompt, then make small single-change edits.
- Re-specify critical constraints when you iterate.
- Prefer one targeted follow-up at a time over rewriting the whole prompt.
Fallback-only execution controls
quality,input_fidelity, explicit masks, output format, and output paths are fallback-only execution controls.- Do not assume they are built-in
image_gentool arguments. - If the user explicitly chooses CLI fallback, see
references/cli.mdandreferences/image-api.mdfor those controls.
Use-case tips
Generate:
- photorealistic-natural: Prompt as if a real photo is captured in the moment; use photography language (lens, lighting, framing); call for real texture; avoid over-stylized polish unless requested.
- product-mockup: Describe the product/packaging and materials; ensure clean silhouette and label clarity; if in-image text is needed, require verbatim rendering and specify typography.
- ui-mockup: Describe the target fidelity first (shippable mockup or low-fi wireframe), then focus on layout, hierarchy, and practical UI elements; avoid concept-art language.
- infographic-diagram: Define the audience and layout flow; label parts explicitly; require verbatim text.
- logo-brand: Keep it simple and scalable; ask for a strong silhouette and balanced negative space; avoid decorative flourishes unless requested.
- illustration-story: Define panels or scene beats; keep each action concrete.
- stylized-concept: Specify style cues, material finish, and rendering approach (3D, painterly, clay) without inventing new story elements.
- historical-scene: State the location/date and required period accuracy; constrain clothing, props, and environment to match the era.
Edit:
- text-localization: Change only the text; preserve layout, typography, spacing, and hierarchy; no extra words or reflow unless needed.
- identity-preserve: Lock identity (face, body, pose, hair, expression); change only the specified elements; match lighting and shadows.
- precise-object-edit: Specify exactly what to remove/replace; preserve surrounding texture and lighting; keep everything else unchanged.
- lighting-weather: Change only environmental conditions (light, shadows, atmosphere, precipitation); keep geometry, framing, and subject identity.
- background-extraction: Request a clean cutout; crisp silhouette; no halos; preserve label text exactly; no restyling.
- style-transfer: Specify style cues to preserve (palette, texture, brushwork) and what must change; add
no extra elementsto prevent drift. - compositing: Reference inputs by index; specify what moves where; match lighting, perspective, and scale; keep the base framing unchanged.
- sketch-to-render: Preserve layout, proportions, and perspective; choose materials and lighting that support the supplied sketch without adding new elements.
Where to find copy/paste recipes
For copy/paste prompt specs (examples only), see references/sample-prompts.md. This file focuses on principles, specificity, and iteration patterns.
references/sample-prompts.md (verbatim)
Sample prompts (copy/paste)
These prompt recipes are shared across both top-level modes of the skill:
- built-in
image_gentool (default) - explicit
scripts/image_gen.pyCLI fallback
Use these as starting points. They are intentionally complete prompt recipes, not the default amount of augmentation to add to every user request.
When adapting a user's prompt:
- keep user-provided requirements
- only add detail according to the specificity policy in
SKILL.md - do not treat every example below as permission to invent extra story elements
The labeled lines are prompt scaffolding, not a closed schema. Asset type and Input images are prompt-only scaffolding; the CLI does not expose them as dedicated flags.
Execution details such as explicit CLI flags, quality, input_fidelity, masks, output formats, and local output paths depend on mode. Use the built-in tool by default; only apply CLI-specific controls after the user explicitly opts into fallback mode.
For prompting principles (structure, specificity, invariants, iteration), see references/prompting.md.
Generate
photorealistic-natural
Use case: photorealistic-natural
Primary request: candid photo of an elderly sailor on a small fishing boat adjusting a net
Scene/backdrop: coastal water with soft haze
Subject: weathered skin with wrinkles and sun texture
Style/medium: photorealistic candid photo
Composition/framing: medium close-up, eye-level
Lighting/mood: soft coastal daylight, shallow depth of field, subtle film grain
Materials/textures: real skin texture, worn fabric, salt-worn wood
Constraints: natural color balance; no heavy retouching; no glamorization; no watermark
Avoid: studio polish; staged look
product-mockup
Use case: product-mockup
Primary request: premium product photo of a matte black shampoo bottle with a minimal label
Scene/backdrop: clean studio gradient from light gray to white
Subject: single bottle centered with subtle reflection
Style/medium: premium product photography
Composition/framing: centered, slight three-quarter angle, generous padding
Lighting/mood: softbox lighting, clean highlights, controlled shadows
Materials/textures: matte plastic, crisp label printing
Constraints: no logos or trademarks; no watermark
ui-mockup
Use case: ui-mockup
Primary request: mobile app home screen for a local farmers market with vendors and daily specials
Asset type: mobile app screen
Style/medium: realistic product UI, not concept art
Composition/framing: clean vertical mobile layout with clear hierarchy
Constraints: practical layout, clear typography, no logos or trademarks, no watermark
infographic-diagram
Use case: infographic-diagram
Primary request: detailed infographic of an automatic coffee machine flow
Scene/backdrop: clean, light neutral background
Subject: bean hopper -> grinder -> brew group -> boiler -> water tank -> drip tray
Style/medium: clean vector-like infographic with clear callouts and arrows
Composition/framing: vertical poster layout, top-to-bottom flow
Text (verbatim): "Bean Hopper", "Grinder", "Brew Group", "Boiler", "Water Tank", "Drip Tray"
Constraints: clear labels, strong contrast, no logos or trademarks, no watermark
logo-brand
Use case: logo-brand
Primary request: original logo for "Field & Flour", a local bakery
Style/medium: vector logo mark; flat colors; minimal
Composition/framing: single centered logo on a plain background with generous padding
Constraints: strong silhouette, balanced negative space; original design only; no gradients unless essential; no trademarks; no watermark
illustration-story
Use case: illustration-story
Primary request: 4-panel comic about a pet left alone at home
Scene/backdrop: cozy living room across panels
Subject: pet reacting to the owner leaving, then relaxing, then returning to a composed pose
Style/medium: comic illustration with clear panels
Composition/framing: 4 equal-sized vertical panels, readable actions per panel
Constraints: no text; no logos or trademarks; no watermark
stylized-concept
Use case: stylized-concept
Primary request: cavernous hangar interior with tall support beams and drifting fog
Scene/backdrop: industrial hangar interior, deep scale, light haze
Subject: compact shuttle parked near the center
Style/medium: cinematic concept art, industrial realism
Composition/framing: wide-angle, low-angle
Lighting/mood: volumetric light rays cutting through fog
Constraints: no logos or trademarks; no watermark
historical-scene
Use case: historical-scene
Primary request: outdoor crowd scene in Bethel, New York on August 16, 1969
Scene/backdrop: open field with period-appropriate staging
Subject: crowd in period-accurate clothing, authentic environment
Style/medium: photorealistic photo
Composition/framing: wide shot, eye-level
Constraints: period-accurate details; no modern objects; no logos or trademarks; no watermark
Asset type templates (taxonomy-aligned)
Website assets template
Use case: <photorealistic-natural|stylized-concept|product-mockup|infographic-diagram|ui-mockup>
Asset type: <hero image / section illustration / blog header>
Primary request: <short description>
Scene/backdrop: <environment or abstract backdrop>
Subject: <main subject>
Style/medium: <photo/illustration/3D>
Composition/framing: <wide/centered; note usable negative space only if needed>
Lighting/mood: <soft/bright/neutral>
Color palette: <brand colors or neutral>
Constraints: <no text; no logos; no watermark; leave room for UI if needed>
Website assets example: minimal hero background
Use case: stylized-concept
Asset type: landing page hero background
Primary request: minimal abstract background with a soft gradient and subtle texture
Style/medium: matte illustration / soft-rendered abstract background
Composition/framing: wide composition with usable negative space for page copy
Lighting/mood: gentle studio glow
Color palette: restrained neutral palette
Constraints: no text; no logos; no watermark
Website assets example: feature section illustration
Use case: stylized-concept
Asset type: feature section illustration
Primary request: simple abstract shapes suggesting connection and flow
Scene/backdrop: subtle light-gray backdrop with faint texture
Style/medium: flat illustration; soft shadows; restrained contrast
Composition/framing: centered cluster; open margins for UI
Color palette: muted neutral palette
Constraints: no text; no logos; no watermark
Website assets example: blog header image
Use case: photorealistic-natural
Asset type: blog header image
Primary request: overhead desk scene with notebook, pen, and coffee cup
Scene/backdrop: warm wooden tabletop
Style/medium: photorealistic photo
Composition/framing: wide crop with clean room for page copy
Lighting/mood: soft morning light
Constraints: no text; no logos; no watermark
Game assets template
Use case: stylized-concept
Asset type: <game environment concept art / game character concept / game UI icon / tileable game texture>
Primary request: <biome/scene/character/icon/material>
Scene/backdrop: <location + set dressing> (if applicable)
Subject: <main focal element(s)>
Style/medium: <realistic/stylized>; <concept art / character render / UI icon / texture>
Composition/framing: <wide/establishing/top-down>; <camera angle>; <focal point placement>
Lighting/mood: <time of day>; <mood>; <volumetric/fog/etc>
Constraints: no logos or trademarks; no watermark
Game assets example: environment concept art
Use case: stylized-concept
Asset type: game environment concept art
Primary request: cavernous hangar interior with tall support beams and drifting fog
Scene/backdrop: industrial hangar interior, deep scale, light haze
Subject: compact shuttle parked near the center
Style/medium: cinematic concept art, industrial realism
Composition/framing: wide-angle, low-angle
Lighting/mood: volumetric light rays cutting through fog
Constraints: no logos or trademarks; no watermark
Game assets example: character concept
Use case: stylized-concept
Asset type: game character concept
Primary request: desert scout character with layered travel gear
Subject: long coat, satchel, practical travel clothing
Style/medium: character render; stylized realism
Composition/framing: neutral hero pose on a simple backdrop
Constraints: no logos or trademarks; no watermark
Game assets example: UI icon
Use case: stylized-concept
Asset type: game UI icon
Primary request: round shield icon with a subtle rune pattern
Style/medium: painted game UI icon
Composition/framing: centered icon; generous padding; clear silhouette
Constraints: no text; no background scene elements; no logos or trademarks; no watermark
Game assets example: tileable texture
Use case: stylized-concept
Asset type: tileable game texture
Primary request: worn sandstone blocks
Style/medium: seamless tileable texture; PBR-ish look
Scene/backdrop: neutral lighting reference only
Constraints: seamless edges; no obvious focal elements; no text; no logos or trademarks; no watermark
Wireframe template
Use case: ui-mockup
Asset type: website wireframe
Primary request: <page or flow to sketch>
Style/medium: low-fi grayscale wireframe
Composition/framing: <landscape or portrait to match expected device>
Subject: <sections in order; grid/columns; key labels>
Constraints: no color; no logos; no real photos; no watermark
Wireframe example: homepage (desktop)
Use case: ui-mockup
Asset type: website wireframe
Primary request: SaaS homepage layout with clear hierarchy
Style/medium: low-fi grayscale wireframe
Subject: top nav; hero with headline and CTA; three feature cards; testimonial strip; pricing preview; footer
Composition/framing: landscape desktop layout
Constraints: label major blocks; no color; no logos; no real photos; no watermark
Wireframe example: pricing page
Use case: ui-mockup
Asset type: website wireframe
Primary request: pricing page layout with comparison table
Style/medium: low-fi grayscale wireframe
Subject: header; plan toggle; 3 pricing cards; comparison table; FAQ accordion; footer
Composition/framing: desktop or tablet layout
Constraints: label key areas; no color; no logos; no real photos; no watermark
Wireframe example: mobile onboarding flow
Use case: ui-mockup
Asset type: mobile onboarding wireframe
Primary request: three-screen mobile onboarding flow
Style/medium: low-fi grayscale wireframe
Subject: screen 1 headline and CTA; screen 2 feature bullets; screen 3 form fields and CTA
Composition/framing: portrait mobile layout
Constraints: label screens and blocks; no color; no logos; no real photos; no watermark
Logo template
Use case: logo-brand
Asset type: logo concept
Primary request: <brand idea or symbol concept>
Style/medium: vector logo mark; flat colors; minimal
Composition/framing: centered mark; clear silhouette; generous margin
Color palette: <1-2 colors; high contrast>
Text (verbatim): "<exact name>" (only if needed)
Constraints: no gradients; no mockups; no 3D; no watermark
Logo example: abstract symbol mark
Use case: logo-brand
Asset type: logo concept
Primary request: geometric leaf symbol suggesting sustainability and growth
Style/medium: vector logo mark; flat colors; minimal
Composition/framing: centered mark; clear silhouette
Color palette: deep green and off-white
Constraints: no text unless requested; no gradients; no mockups; no 3D; no watermark
Logo example: monogram mark
Use case: logo-brand
Asset type: logo concept
Primary request: interlocking monogram of the letters "AV"
Style/medium: vector logo mark; flat colors; minimal
Composition/framing: centered mark; balanced spacing
Color palette: black on white
Constraints: no gradients; no mockups; no 3D; no watermark
Logo example: wordmark
Use case: logo-brand
Asset type: logo concept
Primary request: clean wordmark for a modern studio
Style/medium: vector wordmark; flat colors; minimal
Text (verbatim): "Studio North"
Composition/framing: centered text; even letter spacing
Constraints: no gradients; no mockups; no 3D; no watermark
Edit
text-localization
Use case: text-localization
Input images: Image 1: original infographic
Primary request: replace "Bean Hopper", "Grinder", "Brew Group", "Boiler", "Water Tank", and "Drip Tray" with "Tolva", "Molino", "Grupo de infusión", "Caldera", "Depósito de agua", and "Bandeja de goteo"
Constraints: change only the text; preserve layout, typography, spacing, and hierarchy; no extra words; do not alter logos or imagery
identity-preserve
Use case: identity-preserve
Input images: Image 1: person photo; Image 2..N: clothing references
Primary request: replace only the clothing with the provided garments
Constraints: preserve face, body shape, pose, hair, expression, and identity; match lighting and shadows; keep the background unchanged; no accessories or text
precise-object-edit
Use case: precise-object-edit
Input images: Image 1: room photo
Primary request: replace only the white chairs with wooden chairs
Constraints: preserve camera angle, room lighting, floor shadows, and surrounding objects; keep all other aspects unchanged
lighting-weather
Use case: lighting-weather
Input images: Image 1: original photo
Primary request: make it look like a winter evening with gentle snowfall
Constraints: preserve subject identity, geometry, camera angle, and composition; change only lighting, atmosphere, and weather
background-extraction
Use case: background-extraction
Input images: Image 1: product photo
Primary request: isolate the product on a clean transparent background
Constraints: crisp silhouette; no halos or fringing; preserve label text exactly; no restyling
style-transfer
Use case: style-transfer
Input images: Image 1: style reference
Primary request: apply Image 1's visual style to a man riding a motorcycle on a plain white backdrop
Constraints: preserve palette, texture, and brushwork; no extra elements
compositing
Use case: compositing
Input images: Image 1: base scene; Image 2: subject to insert
Primary request: place the subject from Image 2 next to the person in Image 1
Constraints: match lighting, perspective, and scale; keep the base framing unchanged; no extra elements
sketch-to-render
Use case: sketch-to-render
Input images: Image 1: drawing
Primary request: turn the drawing into a photorealistic image
Constraints: preserve layout, proportions, and perspective; choose realistic materials and lighting; do not add new elements or text
Back to openai/skills (Skills Catalog for Codex) or Agent skills.