cli-creator skill (openai/skills)

From Public Agent Wiki

What it does. Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script. Use when the user wants Codex to create a command-line tool that can run from any repo, expose composable read/write commands, return stable JSON, manage auth, and pair with a companion skill. Part of openai/skills (Skills Catalog for Codex) (openai/skills).

Upstream openai/skills
Skill file skills/.curated/cli-creator/SKILL.md
License Apache-2.0 (skill folder LICENSE.txt)
Author OpenAI
Fetched 2026-09-10

Install

  • Codex: $skill-installer installs from this catalog ($cli-creator invokes it); other agents: npx skills add openai/skills --skill cli-creator.
  • Raw file: curl -sL https://raw.githubusercontent.com/openai/skills/HEAD/skills/.curated/cli-creator/SKILL.md

SKILL.md (verbatim)

name: cli-creator
description: Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script. Use when the user wants Codex to create a command-line tool that can run from any repo, expose composable read/write commands, return stable JSON, manage auth, and pair with a companion skill.

CLI Creator

Create a real CLI that future Codex threads can run by command name from any working directory.

This skill is for durable tools, not one-off scripts. If a short script in the current repo solves the task, write the script there instead.

Start

Name the target tool, its source, and the first real jobs it should do:

  • Source: API docs, OpenAPI JSON, SDK docs, curl examples, browser app, existing internal script, article, or working shell history.
  • Jobs: literal reads/writes such as list drafts, download failed job logs, search messages, upload media, read queue schedule.
  • Install name: a short binary name such as ci-logs, slack-cli, sentry-cli, or buildkite-logs.

Prefer a new folder under ~/code/clis/<tool-name> when the user wants a personal tool and has not named a repo.

Before scaffolding, check whether the proposed command already exists:

command -v <tool-name> || true

If it exists, choose a clearer install name or ask the user.

Choose the Runtime

Before choosing, inspect the user's machine and source material:

command -v cargo rustc node pnpm npm python3 uv || true

Then choose the least surprising toolchain:

  • Default to Rust for a durable CLI Codex should run from any repo: one fast binary, strong argument parsing, good JSON handling, easy copy/install into ~/.local/bin.
  • Use TypeScript/Node when the official SDK, auth helper, browser automation library, or existing repo tooling is the reason the CLI can be better.
  • Use Python when the source is data science, local file transforms, notebooks, SQLite/CSV/JSON analysis, or Python-heavy admin tooling that can still be installed as a durable command.

Do not pick a language that adds setup friction unless it materially improves the CLI. If the best language is not installed, either install the missing toolchain with the user's approval or choose the next-best installed option.

State the choice in one sentence before scaffolding, including the reason and the installed toolchain you found.

Command Contract

Sketch the command surface in chat before coding. Include the binary name, discovery commands, resolve or ID-lookup commands, read commands, write commands, raw escape hatch, auth/config choice, and PATH/install command.

When designing the command surface, read references/agent-cli-patterns.md for the expected composable CLI shape.

Build toward this surface:

  • tool-name --help shows every major capability.
  • tool-name --json doctor verifies config, auth, version, endpoint reachability, and missing setup.
  • tool-name init ... stores local config when env-only auth is painful.
  • Discovery commands find accounts, projects, workspaces, teams, queues, channels, repos, dashboards, or other top-level containers.
  • Resolve commands turn names, URLs, slugs, permalinks, customer input, or build links into stable IDs so future commands do not repeat broad searches.
  • Read commands fetch exact objects and list/search collections. Paginated lists support a bounded --limit, cursor, offset, or clearly documented default.
  • Write commands do one named action each: create, update, delete, upload, schedule, retry, comment, draft. They accept the narrowest stable resource ID, support --dry-run, draft, or preview first when the service allows it, and do not hide writes inside broad commands such as fix, debug, or auto.
  • --json returns stable machine-readable output.
  • A raw escape hatch exists: request, tool-call, api, or the nearest honest name.

Do not expose only a generic request command. Give Codex high-level verbs for the repeated jobs.

Document the JSON policy in the CLI README or equivalent: API pass-through versus CLI envelope, success shape, error shape, and one example for each command family. Under --json, errors must be machine-readable and must not contain credentials.

Auth and Config

Support the boring paths first, in this precedence order:

  1. Environment variable using the service's standard name, such as GITHUB_TOKEN.
  2. User config under ~/.<tool-name>/config.toml or another simple documented path.
  3. --api-key or a tool-specific token flag only for explicit one-off tests. Prefer env/config for normal use because flags can leak into shell history or process listings.

Never print full tokens. doctor --json should say whether a token is available, the auth source category (flag, env, config, provider default, or missing), and what setup step is missing.

If the CLI can run without network or auth, make that explicit in doctor --json: report fixture/offline mode, whether fixture data was found, and whether auth is not required for that mode.

For internal web apps sourced from DevTools curls, create sanitized endpoint notes before implementing: resource name, method/path, required headers, auth mechanism, CSRF behavior, request body, response ID fields, pagination, errors, and one redacted sample response. Never commit copied cookies, bearer tokens, customer secrets, or full production payloads.

Use screenshots to infer workflow, UI vocabulary, fields, and confirmation points. Do not treat screenshots as API evidence unless they are paired with a network request, export, docs page, or fixture.

Build Workflow

  1. Read the source just enough to inventory resources, auth, pagination, IDs, media/file flows, rate limits, and dangerous write actions. If the docs expose OpenAPI, download or inspect it before naming commands.
  2. Sketch the command list in chat. Keep names short and shell-friendly.
  3. Scaffold the CLI with a README or equivalent repo-facing instructions.
  4. Implement doctor, discovery, resolve, read commands, one narrow draft or dry-run write path if requested, and the raw escape hatch.
  5. Install the CLI on PATH so tool-name ... works outside the source folder.
  6. Smoke test from another repo or /tmp, not only with cargo run or package-manager wrappers. Run command -v <tool-name>, <tool-name> --help, and <tool-name> --json doctor.
  7. Run format, typecheck/build, unit tests for request builders, pagination/request-body builders, no-auth doctor, help output, and at least one fixture, dry-run, or live read-only API call.

If a live write is needed for confidence, ask first and make it reversible or draft-only.

When the source is an existing script or shell history, split the working invocation into real phases: setup, discovery, download/export, transform/index, draft, upload, poll, live write. Preserve the flags, paths, and environment variables the user already relies on, then wrap the repeatable phases with stable IDs, bounded JSON, and file outputs.

For raw escape hatches, support read-only calls first. Do not run raw non-GET/HEAD requests against a live service unless the user asked for that specific write.

For media, artifact, or presigned upload flows, test each phase separately: create upload, transfer bytes, poll/read processing status, then attach or reference the resulting ID.

For fixture-backed prototypes, keep fixtures in a predictable project path and make the CLI locate them after installation. Smoke-test from /tmp to catch binaries that only work inside the source folder.

For log-oriented CLIs, keep deterministic snippet extraction separate from model interpretation. Prefer a command that emits filenames, line numbers or byte ranges, matched rules, and short excerpts.

Rust Defaults

When building in Rust, use established crates instead of custom parsers:

  • clap for commands and help
  • reqwest for HTTP
  • serde / serde_json for payloads
  • toml for small config files
  • anyhow for CLI-shaped error context

Add a Makefile target such as make install-local that builds release and installs the binary into ~/.local/bin.

TypeScript/Node Defaults

When building in TypeScript/Node, keep the CLI installable as a normal command:

  • commander or cac for commands and help
  • native fetch, the official SDK, or the user's existing HTTP helper for API calls
  • zod only where external payload validation prevents real breakage
  • package.json bin entry for the installed command
  • tsup, tsx, or tsc using the repo's existing convention

Add an install path such as pnpm install, pnpm build, and pnpm link --global, or a Makefile target that installs a small wrapper into ~/.local/bin.

Python Defaults

When building in Python, prefer boring standard-library pieces unless the workflow needs more:

  • argparse for commands and help, or typer when subcommands would otherwise get messy
  • urllib.request / urllib.parse, requests, or httpx for HTTP, matching what is already installed or already used nearby
  • json, csv, sqlite3, pathlib, and subprocess for local files, exports, databases, and existing scripts
  • pyproject.toml console script or a small executable wrapper for the installed command
  • uv or a virtualenv only when dependencies are actually needed

Add a Makefile target such as make install-local that installs the command on PATH and document whether it depends on uv, a virtualenv, or only system Python.

Companion Skill

After the CLI works, create or update a small skill for it. Use $skill-creator when it is available. Use $CODEX_HOME/skills/<tool-name>/SKILL.md for a personal companion skill unless the user names a repo-local .codex/skills/... path or another skill repo.

Write the companion skill in the order a future Codex thread should use the CLI, not as a tour of every feature. Explain:

  • How to verify the installed command exists.
  • Which command to run first.
  • How auth is configured.
  • Which discovery command finds the common ID.
  • The safe read path.
  • The intended draft/write path.
  • The raw escape hatch.
  • What not to do without explicit user approval.
  • Three copy-pasteable command examples.

Keep API reference details in the CLI docs or a skill reference file. Keep the skill focused on ordering, safety, and examples future Codex threads should actually run.

Other files in this skill

references/agent-cli-patterns.md (verbatim)

Codex CLI Patterns

Use this reference when designing the command surface for a new CLI Codex should run.

Mental model

The CLI is Codex's command layer. It should turn a service, app, API, log source, or database into shell commands Codex can run repeatedly from any repo.

Good CLIs for Codex expose composable primitives. Avoid a single command that tries to "do the whole investigation" when smaller discover, read, resolve, download, inspect, draft, and upload commands would compose better.

Help is interface

Write --help for a future Codex thread that only has the binary and a vague task. Each command should have a short description and flags with literal names from the product or API.

Good top-level help should answer:

  • What containers can I discover?
  • What exact objects can I read?
  • What stable IDs can I resolve?
  • What files can I download or upload?
  • Which write actions exist?
  • What is the raw escape hatch?

Prefer this command shape

Use product nouns, then verbs:

tool-name --json doctor
tool-name --json accounts list
tool-name --json projects list
tool-name --json channels resolve --name codex
tool-name --json messages search "exact phrase"
tool-name --json messages context <message-id> --before 3 --after 3
tool-name --json logs download <build-url> --failed --out ./logs
tool-name --json media upload --file ./image.png
tool-name --json drafts create --body-file draft.json

For APIs whose native noun is already strong, direct verbs can be fine:

tool-name --json social-sets
tool-name --json drafts list --social-set <id>
tool-name --json request get /v2/me

The important rule is consistency. Do not mix many styles unless the product vocabulary demands it.

Useful shapes from mature CLIs

Prefer these patterns over clever agent-only abstractions:

# Field-selected structured output: make common reads scriptable.
tool-name issues list --json number,title,url,state
tool-name issues list --json number,title --jq '.[] | select(.state == "open")'

# Human text by default, full API object when requested.
tool-name pods get <name>
tool-name pods get <name> -o json

# Product workflow commands, not just REST nouns.
tool-name logs tail
tool-name webhooks listen --forward-to localhost:4242/webhooks
tool-name webhooks trigger checkout.completed

Only implement filtering or templating if the user will actually need it. Stable JSON plus narrow read commands are the baseline.

Discovery, resolve, read, context

Design first-pass commands in this order:

  1. Discover broad containers: workspaces, accounts, social sets, repos, projects, channels, queues.
  2. Resolve human input into IDs: user names, channel names, permalinks, PR URLs, build URLs, customer slugs.
  3. Read an exact object: issue, event, thread, draft, customer, job, run, media item.
  4. Context around an anchor when useful: nearby messages, parent thread, surrounding logs, audit history.

Do not force Codex to repeatedly search when it already has a stable ID.

Text, JSON, files, exit codes

Support human text by default if it helps. Support --json everywhere Codex will parse or pipe results.

For --json:

  • Emit JSON to stdout only.
  • Send progress and diagnostics to stderr.
  • Keep success and error shapes documented.
  • Redact tokens, cookies, customer secrets, private headers, and unrelated payloads.

For downloads and exports:

  • Write files under a user-provided --out path when possible.
  • In JSON output, return the file path, byte count if cheap, source URL or ID, and follow-up command.

For exit codes:

  • Exit zero when the command succeeded, including an empty result.
  • Exit nonzero for auth failure, invalid input, network failure, parse failure, API error, or incomplete upload/download.
  • Make doctor --json usable even when auth is missing. It should report missing auth rather than crashing.

Pagination and breadth

Start shallow by default. Add explicit knobs for breadth:

tool-name --json messages search "topic" --limit 10
tool-name --json messages search "topic" --limit 50 --all-pages --max-pages 3
tool-name --json drafts list --limit 20 --offset 40

Return next_cursor, next_url, offset, page_count, or whatever is real for the provider.

Raw escape hatch

The raw command is a repair hatch, not the main interface.

Good raw commands still use configured auth, base URL, JSON parsing, redaction, status/error handling, and --json.

Make reads easy:

tool-name --json request get /v2/me

Treat raw writes as live writes. Do not hide POST/PUT/PATCH/DELETE behind a "debug" command.

Companion skill pattern

The companion skill should be smaller than the CLI README. It should teach the path through the tool:

Start with:

tool-name --json doctor
tool-name --json accounts list

For [common job]:

tool-name --json ... 
tool-name --json ...

Rules:

- Prefer installed `tool-name` on PATH.
- Use --json when analyzing output.
- Create drafts by default.
- Do not publish/delete/retry/submit unless the user asked.
- Use `request get ...` only when high-level commands are missing.

Include JSON shape notes only when Codex needs them to choose the next command.

Back to openai/skills (Skills Catalog for Codex) or Agent skills.