retro skill (gstack) (part 2)

From Public Agent Wiki

Part 2 of 2 of retro skill (gstack) (retro/SKILL.md in garrytan/gstack); the SKILL.md text continues verbatim from the previous part.

SKILL.md (verbatim, continued)

  }

Step 14: Write the Narrative

STOP. Before writing the retrospective narrative (Step 14, after all metrics are computed and compared), Read ~/.claude/skills/gstack/retro/sections/report-format.md and execute it in full. Do not work from memory โ€” that section is the source of truth for this step.

After delivering the repo-scoped report, run the following learning capture and result-save steps, then stop. Do not fall through into Global Retrospective Mode.

Capture Learnings

If you discovered a non-obvious pattern, pitfall, or architectural insight during this session, log it for future sessions:

~/.claude/skills/gstack/bin/gstack-learnings-log '{"skill":"retro","type":"TYPE","key":"SHORT_KEY","insight":"DESCRIPTION","confidence":N,"source":"SOURCE","files":["path/to/relevant/file"]}'

Types: pattern (reusable approach), pitfall (what NOT to do), preference (user stated), architecture (structural decision), tool (library/framework insight), operational (project environment/CLI/workflow knowledge).

Sources: observed (you found this in the code), user-stated (user told you), inferred (AI deduction), cross-model (both Claude and Codex agree).

Confidence: 1-10. Be honest. An observed pattern you verified in the code is 8-9. An inference you're not sure about is 4-5. A user preference they explicitly stated is 10.

files: Include the specific file paths this learning references. This enables staleness detection: if those files are later deleted, the learning can be flagged.

Only log genuine discoveries. Don't log obvious things. Don't log things the user already knows. A good test: would this insight save time in a future session? If yes, log it.


Global Retrospective Mode

/retro global [window] follows only this flow and works outside a git repo.

Global Step 1: Compute time window

Same midnight-aligned logic as the regular retro. Default 7d. The second argument after global is the window (e.g., 14d, 30d, 24h).

Global Step 2: Run discovery

Locate and run the discovery script using this fallback chain:

DISCOVER_BIN=""
[ -x ~/.claude/skills/gstack/bin/gstack-global-discover ] && DISCOVER_BIN=~/.claude/skills/gstack/bin/gstack-global-discover
[ -z "$DISCOVER_BIN" ] && [ -x .claude/skills/gstack/bin/gstack-global-discover ] && DISCOVER_BIN=.claude/skills/gstack/bin/gstack-global-discover
[ -z "$DISCOVER_BIN" ] && which gstack-global-discover >/dev/null 2>&1 && DISCOVER_BIN=$(which gstack-global-discover)
[ -z "$DISCOVER_BIN" ] && [ -f bin/gstack-global-discover.ts ] && DISCOVER_BIN="bun run bin/gstack-global-discover.ts"
echo "DISCOVER_BIN: $DISCOVER_BIN"

If no binary is found, tell the user: "Discovery script not found. Run bun run build in the gstack directory to compile it." and stop.

Run the discovery:

$DISCOVER_BIN --since "<window>" --format json 2>/tmp/gstack-discover-stderr

Read the stderr output from /tmp/gstack-discover-stderr for diagnostic info. Parse the JSON output from stdout.

If total_sessions is 0, say: "No AI coding sessions found in the last <window>. Try a longer window: /retro global 30d" and stop.

Global Step 3: Run git log on each discovered repo

For each repo in the discovery JSON's repos array, find the first valid path in paths[] (directory exists with .git/). If no valid path exists, skip the repo and note it.

For local-only repos (where remote starts with local:): skip git fetch and use the local default branch. Use git log HEAD instead of git log origin/$DEFAULT.

For repos with remotes:

git -C <path> fetch origin --quiet 2>/dev/null

Detect the default branch for each repo: first try git symbolic-ref refs/remotes/origin/HEAD, then check common branch names (main, master), then fall back to git rev-parse --abbrev-ref HEAD. Use the detected branch as <default> in the commands below.

# Commits with stats
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%H|%aN|%ai|%s" --shortstat

# Commit timestamps for session detection, streak, and context switching
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%at|%aN|%ai|%s" | sort -n

# Per-author commit counts
git -C <path> shortlog origin/$DEFAULT --since="<start_date>T00:00:00" -sn --no-merges

# PR/MR numbers from commit messages (GitHub #NNN, GitLab !NNN)
git -C <path> log origin/$DEFAULT --since="<start_date>T00:00:00" --format="%s" | grep -oE '[#!][0-9]+' | sort -t'#' -k1 | uniq

For repos that fail (deleted paths, network errors): skip and note "N repos could not be reached."

Global Step 4: Compute global shipping streak

For each repo, get commit dates (capped at 365 days):

git -C <path> log origin/$DEFAULT --since="365 days ago" --format="%ad" --date=format:"%Y-%m-%d" | sort -u

Union all dates across all repos. Count backward from today โ€” how many consecutive days have at least one commit to ANY repo? If the streak hits 365 days, display as "365+ days".

Global Step 5: Compute context switching metric

From the commit timestamps gathered in Step 3, group by date. For each date, count how many distinct repos had commits that day. Report:

  • Average repos/day
  • Maximum repos/day
  • Which days were focused (1 repo) vs. fragmented (3+ repos)

Global Step 6: Per-tool productivity patterns

From the discovery JSON, analyze tool usage patterns:

  • Which AI tool is used for which repos (exclusive vs. shared)
  • Session count per tool
  • Behavioral patterns (e.g., "Codex used exclusively for myapp, Claude Code for everything else")

Global Step 7: Aggregate and draft narrative

Draft the report below without publishing it yet. Load history in Global Step 8, insert its trends table after All Projects Overview, then save the completed snapshot in Global Step 9 and deliver the report. Reuse the drafted tweetable summary in the snapshot.

Output the screenshot-friendly personal card first, then the team/project breakdown.


Tweetable summary (first line, before everything else):

Week of Mar 14: 5 projects, 138 commits, 250k LOC across 5 repos | 48 AI sessions | Streak: 52d ๐Ÿ”ฅ

๐Ÿš€ Your Week: [user name] โ€” [date range]

Filter per-repo data by git config user.name and aggregate personal totals. The card contains only this user's stats, not team totals. Use a left border only; pad names to the longest name and never truncate them.

โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
โ•‘  [USER NAME] โ€” Week of [date]
โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
โ•‘
โ•‘  [N] commits across [M] projects
โ•‘  +[X]k LOC added ยท [Y]k LOC deleted ยท [Z]k net
โ•‘  [N] AI coding sessions (CC: X, Codex: Y, Gemini: Z)
โ•‘  [N]-day shipping streak ๐Ÿ”ฅ
โ•‘
โ•‘  PROJECTS
โ•‘  โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
โ•‘  [repo_name_full]        [N] commits    +[X]k LOC    [solo/team]
โ•‘  [repo_name_full]        [N] commits    +[X]k LOC    [solo/team]
โ•‘  [repo_name_full]        [N] commits    +[X]k LOC    [solo/team]
โ•‘
โ•‘  SHIP OF THE WEEK
โ•‘  [PR title] โ€” [LOC] lines across [N] files
โ•‘
โ•‘  TOP WORK
โ•‘  โ€ข [1-line description of biggest theme]
โ•‘  โ€ข [1-line description of second theme]
โ•‘  โ€ข [1-line description of third theme]
โ•‘
โ•‘  Powered by gstack
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

Rules for the personal card:

  • Only show repos where the user has commits. Skip repos with 0 commits.
  • Sort repos by user's commit count descending.
  • Widen the card to fit full repo names; align columns.
  • For LOC, use "k" formatting for thousands (e.g., "+64.0k" not "+64010").
  • Role: "solo" if user is the only contributor, "team" if others contributed.
  • Ship of the Week: the user's single highest-LOC PR across ALL repos.
  • Top Work: 3 themes synthesized from commit messages, not a list of commits.
  • The card must explain the user's week without surrounding context.
  • Do NOT include team members, project totals, or context switching data here.

Personal streak: Use the user's own commits across all repos (filtered by --author) to compute a personal streak, separate from the team streak.


Global Engineering Retro: [date range]

Full team/project analysis follows the personal card.

All Projects Overview

Metric Value
Projects active N
Total commits (all repos, all contributors) N
Total LOC +N / -N
AI coding sessions N (CC: X, Codex: Y, Gemini: Z)
Active days N
Global shipping streak (any contributor, any repo) N consecutive days
Context switches/day N avg (max: M)

Per-Project Breakdown

For each repo (sorted by commits descending):

  • Repo name (with % of total commits)
  • Commits, LOC, PRs merged, top contributor
  • Key work (inferred from commit messages)
  • AI sessions by tool

Your Contributions (sub-section within each project): For each project, filter by git config user.name and include:

  • Your commits / total commits (with %)
  • Your LOC (+insertions / -deletions)
  • Your key work (inferred from YOUR commit messages only)
  • Your commit type mix (feat/fix/refactor/chore/docs breakdown)
  • Your biggest ship in this repo (highest-LOC commit or PR)

If the user is the only contributor, say "Solo project โ€” all commits are yours." If the user has 0 commits in a repo (team project they didn't touch this period), say "No commits this period โ€” [N] AI sessions only." and skip the breakdown.

Format:

**Your contributions:** 47/244 commits (19%), +4.2k/-0.3k LOC
  Key work: Writer Chat, email blocking, security hardening
  Biggest ship: PR #605 โ€” Writer Chat eats the admin bar (2,457 ins, 46 files)
  Mix: feat(3) fix(2) chore(1)

Cross-Project Patterns

  • Time allocation across projects (% breakdown, use YOUR commits not total)
  • Peak productivity hours aggregated across all repos
  • Focused vs. fragmented days
  • Context switching trends

Tool Usage Analysis

Per-tool breakdown with behavioral patterns:

  • Claude Code: N sessions across M repos โ€” patterns observed
  • Codex: N sessions across M repos โ€” patterns observed
  • Gemini: N sessions across M repos โ€” patterns observed

Ship of the Week (Global)

Highest-impact PR across ALL projects. Identify by LOC and commit messages.

3 Cross-Project Insights

What the global view reveals that no single-repo retro could show.

3 Habits for Next Week

Considering the full cross-project picture.


Global Step 8: Load history & compare

setopt +o nomatch 2>/dev/null || true  # zsh compat
ls -t ~/.gstack/retros/global-*.json 2>/dev/null | head -5

Only compare against a prior retro with the same window value (e.g., 7d vs 7d). If the most recent prior retro has a different window, skip comparison and note: "Prior global retro used a different window โ€” skipping comparison."

If a matching prior retro exists, load it with the Read tool. Show a Trends vs Last Global Retro table with deltas for key metrics: total commits, LOC, sessions, streak, context switches/day.

If no prior global retros exist, append: "First global retro recorded โ€” run again next week to see trends."

Global Step 9: Save snapshot

mkdir -p ~/.gstack/retros

Determine the next unused sequence number for today, using the same session-reminder date as Global Step 1:

setopt +o nomatch 2>/dev/null || true  # zsh compat
today="<today>"
next=1
while [ -e "$HOME/.gstack/retros/global-${today}-${next}.json" ]; do next=$((next + 1)); done

Use the Write tool to save JSON to ~/.gstack/retros/global-${today}-${next}.json:

{
  "type": "global",
  "date": "2026-03-21",
  "window": "7d",
  "projects": [
    {
      "name": "gstack",
      "remote": "<detected from git remote get-url origin, normalized to HTTPS>",
      "commits": 47,
      "insertions": 3200,
      "deletions": 800,
      "sessions": { "claude_code": 15, "codex": 3, "gemini": 0 }
    }
  ],
  "totals": {
    "commits": 182,
    "insertions": 15300,
    "deletions": 4200,
    "projects": 5,
    "active_days": 6,
    "sessions": { "claude_code": 48, "codex": 8, "gemini": 3 },
    "global_streak_days": 52,
    "avg_context_switches_per_day": 2.1
  },
  "tweetable": "Week of Mar 14: 5 projects, 182 commits, 15.3k LOC | CC: 48, Codex: 8, Gemini: 3 | Focus: gstack (58%) | Streak: 52d"
}

Compare Mode

When the user runs /retro compare (or /retro compare 14d):

  1. Run Steps 0.5-1 for the current window (default 7d) using the midnight-aligned start date (same logic as the main retro โ€” e.g., if today is 2026-03-18 and window is 7d, --since "2026-03-11T00:00:00")
  2. Run gstack-retro-metrics a second time for the immediately prior same-length window, using both --since and --until (e.g., for a 7d window starting 2026-03-11: --since "2026-03-04T00:00:00" --until "2026-03-10T23:59:59")
  3. Compute the windowed metrics in Steps 2-10 for each dataset, keeping current and prior values separate. Run Steps 11-11.5 only for the current report: streaks use full history and the shortcut ledger scans the current tree, so neither is a prior-window metric. Apply the freshness guard only to the current window; an inactive prior window is valid comparison data. For hour windows, capture one explicit end timestamp, then subtract the requested hours twice for the two starts. Git includes --until, so use one second before the current start for the prior end to avoid counting the boundary commit twice.
  4. In place of Step 12's saved-history comparison, show a Current vs Prior Period table for commits, logical SLOC, test ratio, sessions, and fix ratio. Show absolute deltas and percentage changes (ratio changes in percentage points); if the prior value is zero, report absolute change and percentage change as N/A. Highlight the biggest improvements and regressions in the Step 14 narrative.
  5. Run Steps 13-14 and the post-report capture for the current window only; do not persist the prior-window metrics. This comparison works on the first run and does not require saved history.

Tone

  • Encouraging but candid, no coddling
  • Specific and concrete โ€” always anchor in actual commits/code
  • Skip generic praise ("great job!") โ€” say exactly what was good and why
  • Frame improvements as leveling up, not criticism
  • Praise should feel like something you'd actually say in a 1:1 โ€” specific, earned, genuine
  • Growth suggestions should feel like investment advice โ€” "this is worth your time because..." not "you failed at..."
  • Never compare teammates against each other negatively. Each person's section stands on its own.
  • Keep total output around 3000-4500 words (slightly longer to accommodate team sections)
  • Use markdown tables and code blocks for data, prose for narrative
  • Output directly to the conversation โ€” do NOT write to filesystem (except the .context/retros/ JSON snapshot)

Important Rules

  • ALL narrative output goes directly to the user in the conversation. The ONLY file written is the .context/retros/ JSON snapshot.
  • The metrics script analyzes origin/<default> (not local main which may be stale); when RETRO_REF says otherwise, disclose it
  • Display all timestamps in the user's local timezone (do not override TZ)
  • If COMMITS: 0, say so and suggest a different window
  • Round LOC/hour to nearest 50 (the script pre-rounds LOC_PER_SESSION_HOUR)
  • Treat merge commits as PR boundaries
  • Do not read CLAUDE.md or unrelated docs โ€” this skill is self-contained; the CHANGELOG and optional inputs explicitly named above are exceptions
  • On first run (no prior retros), skip saved-history comparisons gracefully; explicit compare mode still computes its prior window
  • Global mode: Does NOT require being inside a git repo. Saves snapshots to ~/.gstack/retros/ (not .context/retros/). Gracefully skip AI tools that aren't installed. Only compare against prior global retros with the same window value. If streak hits 365d cap, display as "365+ days".

Other files in this skill

sections/report-format.md (verbatim)

<!-- AUTO-GENERATED from report-format.md.tmpl โ€” do not edit directly --> <!-- Regenerate: bun run gen:skill-docs -->

Structure the output as:


Tweetable summary (first line, before everything else):

Week of Mar 1: 47 commits (3 contributors), 3.2k LOC, 38% tests, 12 PRs, peak: 10pm | Streak: 47d

Engineering Retro: [date range]

Summary Table

(from Step 2)

(from Step 12, loaded before save โ€” skip if no matching history; in compare mode use Current vs Prior Period from the computed prior window even on the first run)

Time & Session Patterns

(from Steps 3-4)

Narrative interpreting what the team-wide patterns mean:

  • When the most productive hours are and what drives them
  • Whether sessions are getting longer or shorter over time
  • Estimated hours per day of active coding (team aggregate)
  • Notable patterns: do team members code at the same time or in shifts?

Shipping Velocity

(from Steps 5-7)

Narrative covering:

  • Commit type mix and what it reveals
  • PR size distribution and what it reveals about shipping cadence
  • Fix-chain detection (sequences of fix commits on the same subsystem)
  • Version bump discipline

Code Quality Signals

  • Test LOC ratio trend
  • Hotspot analysis (are the same files churning?)
  • Greptile signal ratio and trend (if history exists): "Greptile: X% signal (Y valid catches, Z false positives)"

Test Health

  • Total test files: N (TEST_FILES_TOTAL)
  • Test files changed this period: M (TEST_FILES_CHANGED; not newly added test cases)
  • Regression test commits: list the REGRESSION_COMMIT lines (test(qa):, test(design):, and test: coverage commits)
  • If prior retro exists and has test_health: show delta "Test count: {last} โ†’ {now} (+{delta})"
  • If test ratio < 20%: flag as growth area โ€” "100% test coverage is the goal. Tests make vibe coding safe."

Plan Completion

Check review JSONL logs for plan completion data from /ship runs this period:

setopt +o nomatch 2>/dev/null || true  # zsh compat
eval "$(~/.claude/skills/gstack/bin/gstack-slug 2>/dev/null)"
cat ~/.gstack/projects/$SLUG/*-reviews.jsonl 2>/dev/null | grep '"skill":"ship"' | grep '"plan_items_total"' || echo "NO_PLAN_DATA"

If plan completion data exists within the retro time window:

  • Count branches shipped with plans (entries that have plan_items_total > 0)
  • Compute average completion: sum of plan_items_done / sum of plan_items_total
  • Identify most-skipped item category if data supports it

Output:

Plan Completion This Period:
  {N} branches shipped with plans
  Average completion: {X}% ({done}/{total} items)

If no plan data exists, skip this section silently.

Focus & Highlights

(from Step 8)

  • Focus score with interpretation
  • Ship of the week callout

Shipping Streaks

(from Step 11: team and personal streaks, including broken-streak disclosure)

Shortcut Debt

(from Step 11.5: marker ledger and count, or the clean-ledger statement)

Your Week (personal deep-dive)

(from Step 9, for the current user only)

This is the section the user cares most about. Include:

  • Their personal commit count, LOC, test ratio
  • Their session patterns and peak hours
  • Their focus areas
  • Their biggest ship
  • What you did well (2-3 specific things anchored in commits)
  • Where to level up (1-2 specific, actionable suggestions)

Team Breakdown

(from Step 9, for each teammate โ€” skip if solo repo)

For each teammate (sorted by commits descending), write a section:

[Name]

  • What they shipped: 2-3 sentences on their contributions, areas of focus, and commit patterns
  • Praise: 1-2 specific things they did well, anchored in actual commits. Be genuine โ€” what would you actually say in a 1:1? Examples:
    • "Cleaned up the entire auth module in 3 small, reviewable PRs โ€” textbook decomposition"
    • "Added integration tests for every new endpoint, not just happy paths"
    • "Fixed the N+1 query that was causing 2s load times on the dashboard"
  • Opportunity for growth: 1 specific, constructive suggestion. Frame as investment, not criticism. Examples:
    • "Test coverage on the payment module is at 8% โ€” worth investing in before the next feature lands on top of it"
    • "Most commits land in a single burst โ€” spacing work across the day could reduce context-switching fatigue"
    • "All commits land between 1-4am โ€” sustainable pace matters for code quality long-term"

AI collaboration note: If many commits have Co-Authored-By AI trailers (e.g., Claude, Copilot), note the AI-assisted commit percentage as a team metric. Frame it neutrally โ€” "N% of commits were AI-assisted" โ€” without judgment.

Top 3 Team Wins

Identify the 3 highest-impact things shipped in the window across the whole team. For each:

  • What it was
  • Who shipped it
  • Why it matters (product/architecture impact)

3 Things to Improve

Specific, actionable, anchored in actual commits. Mix personal and team-level suggestions. Phrase as "to get even better, the team could..."

3 Habits for Next Week

Small, practical, realistic. Each must be something that takes <5 minutes to adopt. At least one should be team-oriented (e.g., "review each other's PRs same-day").

(if applicable, from Step 10)

Back to garrytan/gstack (Garry Tan's Claude Code skill suite) or Agent skills.