integrating-sast-into-github-actions-pipeline skill (Anthropic-Cybersecurity-Skills)

From Public Agent Wiki
Contents
  1. Install
  2. SKILL.md (verbatim)
  3. When to Use
  4. Prerequisites
  5. Workflow
  6. Step 1: Configure CodeQL Analysis Workflow
  7. Step 2: Add Semgrep Scanning for Custom Rules
  8. Step 3: Create Custom Semgrep Rules for Organization Patterns
  9. Step 4: Establish Quality Gates with Branch Protection
  10. Step 5: Tune and Suppress False Positives
  11. Step 6: Aggregate and Report Findings
  12. Key Concepts
  13. Tools & Systems
  14. Common Scenarios
  15. Scenario: Monorepo with Multiple Languages Needs Unified SAST
  16. Scenario: Reducing Alert Fatigue from High False Positive Rate
  17. Output Format
  18. Other files in this skill
  19. assets/template.md (verbatim)
  20. GitHub Actions: Combined CodeQL + Semgrep Workflow
  21. CodeQL Custom Configuration
  22. Semgrep Ignore File
  23. Branch Protection Configuration (Terraform)
  24. SARIF Report Aggregation Script
  25. references/api-reference.md (verbatim)
  26. Semgrep CLI
  27. Installation
  28. Scan Commands
  29. JSON Output Structure
  30. Severity Levels
  31. GitHub Actions Integration
  32. Semgrep Action
  33. SARIF Upload
  34. SARIF 2.1.0 Schema
  35. References
  36. references/standards.md (verbatim)
  37. OWASP SAMM - Verification: Security Testing
  38. Maturity Level 1
  39. Maturity Level 2
  40. Maturity Level 3
  41. NIST SSDF (SP 800-218) - Produce Well-Secured Software
  42. PW.7: Review and Analyze Code
  43. PW.8: Test Executable Code
  44. CIS Software Supply Chain Security Guide
  45. Source Code (SC) Controls
  46. Build (BD) Controls
  47. OWASP Top 10 Coverage Matrix
  48. PCI DSS v4.0 Mapping
  49. SOC 2 Trust Service Criteria
  50. references/workflows.md (verbatim)
  51. End-to-End SAST Integration Workflow
  52. CodeQL Analysis Deep Dive
  53. Database Creation Phase
  54. Query Execution Phase
  55. Query Suites
  56. Semgrep Rule Authoring Process
  57. Rule Development Lifecycle
  58. Pattern Operators Reference
  59. SARIF Processing Pipeline
  60. SARIF Structure
  61. Triage and Remediation Workflow
  62. Severity-Based SLA
  63. Finding States

What it does. Integrates CodeQL and Semgrep SAST scanning into GitHub Actions, covering scans on pull requests/pushes, rule tuning to cut false positives, SARIF upload to GitHub Advanced Security, and merge-blocking quality gates for high-severity findings. Use when adding automated code vulnerability detection to CI, enforcing consistent SAST org-wide, or producing SOC 2/PCI DSS/NIST SSDF compliance evidence. Part of mukul975/Anthropic-Cybersecurity-Skills (817 security skills) (mukul975/Anthropic-Cybersecurity-Skills).

Upstream mukul975/Anthropic-Cybersecurity-Skills
Skill file skills/integrating-sast-into-github-actions-pipeline/SKILL.md
License Apache-2.0 (skill folder LICENSE)
Author mukul975
Fetched 2026-09-10

Install

  • npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill integrating-sast-into-github-actions-pipeline, or copy the skill folder into ~/.claude/skills/integrating-sast-into-github-actions-pipeline/.
  • Raw file: curl -sL https://raw.githubusercontent.com/mukul975/Anthropic-Cybersecurity-Skills/HEAD/skills/integrating-sast-into-github-actions-pipeline/SKILL.md

SKILL.md (verbatim)

name: integrating-sast-into-github-actions-pipeline
description: Integrates CodeQL and Semgrep SAST scanning into GitHub Actions, covering scans on pull requests/pushes, rule tuning to cut false positives, SARIF upload to GitHub Advanced Security, and merge-blocking quality gates for high-severity findings. Use when adding automated code vulnerability detection to CI, enforcing consistent SAST org-wide, or producing SOC 2/PCI DSS/NIST SSDF compliance evidence.
domain: cybersecurity
subdomain: devsecops
tags:
- devsecops
- cicd
- sast
- codeql
- semgrep
- secure-sdlc
version: 1.0.0
author: mahipal
license: Apache-2.0
nist_csf:
- PR.PS-01
- GV.SC-07
- ID.IM-04
- PR.PS-04
mitre_attack:
- T1195
- T1554
- T1059.004

Integrating SAST into GitHub Actions Pipeline

When to Use

  • When development teams need automated code-level vulnerability detection on every pull request
  • When security teams require consistent SAST enforcement across all repositories in an organization
  • When migrating from manual or periodic security reviews to continuous security testing
  • When compliance frameworks (SOC 2, PCI DSS, NIST SSDF) require evidence of automated code analysis
  • When multiple languages coexist in a monorepo and need unified scanning under one workflow

Do not use for runtime vulnerability detection (use DAST instead), for scanning third-party dependencies (use SCA tools like Snyk), or for infrastructure-as-code scanning (use Checkov or tfsec).

Prerequisites

  • GitHub repository with GitHub Actions enabled
  • GitHub Advanced Security license (required for CodeQL on private repos; free for public repos)
  • Semgrep account for managed rules and Semgrep App dashboard (free tier available)
  • Repository code in a supported language: Python, JavaScript/TypeScript, Java, C/C++, C#, Go, Ruby, Swift, Kotlin

Workflow

Step 1: Configure CodeQL Analysis Workflow

Create a CodeQL workflow that runs on pull requests and on a weekly schedule to catch vulnerabilities in existing code.

# .github/workflows/codeql-analysis.yml
name: "CodeQL Analysis"

on:
  push:
    branches: [main, develop]
  pull_request:
    branches: [main]
  schedule:
    - cron: '30 2 * * 1'  # Weekly Monday 2:30 AM

jobs:
  analyze:
    name: Analyze (${{ matrix.language }})
    runs-on: ubuntu-latest
    permissions:
      actions: read
      contents: read
      security-events: write

    strategy:
      fail-fast: false
      matrix:
        language: ['javascript', 'python']

    steps:
      - name: Checkout repository
        uses: actions/checkout@v4

      - name: Initialize CodeQL
        uses: github/codeql-action/init@v3
        with:
          languages: ${{ matrix.language }}
          queries: security-extended,security-and-quality

      - name: Autobuild
        uses: github/codeql-action/autobuild@v3

      - name: Perform CodeQL Analysis
        uses: github/codeql-action/analyze@v3
        with:
          category: "/language:${{ matrix.language }}"

Step 2: Add Semgrep Scanning for Custom Rules

Semgrep complements CodeQL with faster scans and support for custom pattern-based rules. Configure it to upload SARIF results to the same GitHub Security tab.

# .github/workflows/semgrep.yml
name: "Semgrep SAST Scan"

on:
  pull_request:
    branches: [main, develop]
  push:
    branches: [main]

jobs:
  semgrep:
    name: Semgrep Scan
    runs-on: ubuntu-latest
    permissions:
      security-events: write
      contents: read

    container:
      image: semgrep/semgrep:latest

    steps:
      - name: Checkout
        uses: actions/checkout@v4

      - name: Run Semgrep
        run: |
          semgrep ci \
            --config auto \
            --config p/owasp-top-ten \
            --config p/cwe-top-25 \
            --sarif --output semgrep-results.sarif \
            --severity ERROR \
            --error
        env:
          SEMGREP_APP_TOKEN: ${{ secrets.SEMGREP_APP_TOKEN }}

      - name: Upload SARIF
        if: always()
        uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: semgrep-results.sarif
          category: semgrep

Step 3: Create Custom Semgrep Rules for Organization Patterns

Write organization-specific rules to catch patterns unique to your codebase, such as deprecated internal APIs or insecure configuration patterns.

# .semgrep/custom-rules.yml
rules:
  - id: hardcoded-database-url
    patterns:
      - pattern: |
          $DB_URL = "...$PROTO://...:...$PASS@..."
    message: |
      Hardcoded database connection string with credentials detected.
      Use environment variables or a secrets manager instead.
    languages: [python, javascript, typescript]
    severity: ERROR
    metadata:
      cwe: "CWE-798: Use of Hard-coded Credentials"
      owasp: "A07:2021 - Identification and Authentication Failures"

  - id: unsafe-deserialization
    patterns:
      - pattern-either:
          - pattern: pickle.loads(...)
          - pattern: yaml.load(..., Loader=yaml.Loader)
          - pattern: yaml.load(..., Loader=yaml.FullLoader)
    message: |
      Unsafe deserialization detected. Use safe alternatives to prevent
      remote code execution vulnerabilities.
    languages: [python]
    severity: ERROR
    metadata:
      cwe: "CWE-502: Deserialization of Untrusted Data"

  - id: missing-csrf-protection
    patterns:
      - pattern: |
          @app.route("...", methods=["POST"])
          def $FUNC(...):
              ...
      - pattern-not-inside: |
          @csrf.exempt
          ...
    message: "POST endpoint may lack CSRF protection."
    languages: [python]
    severity: WARNING

Step 4: Establish Quality Gates with Branch Protection

Configure branch protection rules that require SAST checks to pass before merging, preventing vulnerable code from reaching production branches.

# Use GitHub CLI to set branch protection requiring SAST checks
gh api repos/{owner}/{repo}/branches/main/protection \
  --method PUT \
  --field required_status_checks='{"strict":true,"contexts":["Analyze (javascript)","Analyze (python)","Semgrep Scan"]}' \
  --field enforce_admins=true \
  --field required_pull_request_reviews='{"required_approving_review_count":1}'

Step 5: Tune and Suppress False Positives

Manage false positives through CodeQL query filters and Semgrep nosemgrep annotations to maintain developer trust in scan results.

# codeql-config.yml - Custom CodeQL configuration
name: "Custom CodeQL Config"
queries:
  - uses: security-extended
  - uses: security-and-quality
  - excludes:
      id: js/unused-local-variable
paths-ignore:
  - '**/test/**'
  - '**/tests/**'
  - '**/vendor/**'
  - '**/node_modules/**'
  - '**/*.test.js'
  - '**/*.spec.py'
# Example: Suppressing a known false positive in Semgrep
import subprocess

def run_safe_command(cmd_list):
    # nosemgrep: python.lang.security.audit.dangerous-subprocess-use
    result = subprocess.run(cmd_list, capture_output=True, text=True, shell=False)
    return result.stdout

Step 6: Aggregate and Report Findings

Use the GitHub Security Overview dashboard and configure notifications for security alerts across repositories.

# Query SARIF results via GitHub API for reporting
gh api repos/{owner}/{repo}/code-scanning/alerts \
  --jq '.[] | select(.state=="open") | {rule: .rule.id, severity: .rule.security_severity_level, file: .most_recent_instance.location.path, line: .most_recent_instance.location.start_line}'

# Count open alerts by severity
gh api repos/{owner}/{repo}/code-scanning/alerts \
  --jq '[.[] | select(.state=="open")] | group_by(.rule.security_severity_level) | map({severity: .[0].rule.security_severity_level, count: length})'

Key Concepts

Term Definition
SAST Static Application Security Testing — analyzes source code without executing it to find security vulnerabilities
SARIF Static Analysis Results Interchange Format — standardized JSON format for expressing results from static analysis tools
CodeQL GitHub's semantic code analysis engine that treats code as data and queries it for vulnerability patterns
Semgrep Lightweight static analysis tool using pattern matching to find bugs and security issues across many languages
Security Extended CodeQL query suite that includes additional security queries beyond the default set for deeper analysis
Quality Gate Automated checkpoint that blocks code from progressing through the pipeline unless security criteria are met
False Positive A scan finding that incorrectly identifies secure code as vulnerable, requiring suppression or tuning

Tools & Systems

  • CodeQL: GitHub's semantic code analysis engine with deep dataflow and taint tracking analysis
  • Semgrep: Fast, lightweight pattern-matching SAST tool with 3000+ community rules and custom rule support
  • GitHub Advanced Security: Platform providing code scanning, secret scanning, and dependency review in GitHub
  • SARIF Viewer: VS Code extension for reviewing SARIF results locally during development
  • GitHub Security Overview: Organization-level dashboard aggregating security alerts across all repositories

Common Scenarios

Scenario: Monorepo with Multiple Languages Needs Unified SAST

Context: A platform team manages a monorepo containing Python microservices, TypeScript frontends, and Go infrastructure tools. Security reviews happen manually every quarter, missing vulnerabilities between reviews.

Approach:

  1. Configure CodeQL with a matrix strategy covering Python, JavaScript, and Go languages
  2. Add Semgrep with --config auto to detect language automatically and apply relevant rulesets
  3. Create path-based triggers so only changed language directories trigger their respective scans
  4. Upload all SARIF results to GitHub Security tab with unique categories per tool and language
  5. Set branch protection requiring all SAST jobs to pass before merge
  6. Schedule weekly full-repository scans to catch issues in unchanged code from newly published CVE patterns

Pitfalls: Setting CodeQL to analyze all languages on every PR increases CI time significantly. Use path filters to trigger only relevant language scans. Semgrep's --config auto may enable rules that conflict with CodeQL findings, creating duplicate alerts.

Scenario: Reducing Alert Fatigue from High False Positive Rate

Context: After enabling SAST, developers ignore findings because 40% are false positives, undermining the security program.

Approach:

  1. Export all current alerts and categorize them as true positive, false positive, or informational
  2. Create a custom CodeQL config excluding noisy query IDs that produce the most false positives
  3. Write .semgrepignore patterns for test files, generated code, and vendored dependencies
  4. Establish a weekly triage meeting where security and development leads review new rule additions
  5. Track false positive rate as a metric and target below 15% for developer trust

Pitfalls: Over-suppressing rules to reduce noise can create blind spots. Always validate suppressions against the OWASP Top 10 and CWE Top 25 to ensure critical vulnerability classes remain covered.

Output Format

SAST Pipeline Scan Report
==========================
Repository: org/web-application
Branch: feature/user-auth-refactor
Scan Date: 2026-02-23
Commit: a1b2c3d4

CodeQL Results:
  Language    Queries Run   Findings   Critical   High   Medium
  javascript  312           4          1          2      1
  python      287           2          0          1      1

Semgrep Results:
  Ruleset          Rules Matched   Findings   Errors   Warnings
  auto             1,847           3          1        2
  owasp-top-ten    186             2          1        1
  custom-rules     12              1          0        1

QUALITY GATE: FAILED
  Blocking findings: 2 Critical/High severity issues
  - [CRITICAL] CWE-89: SQL Injection in src/api/users.py:47
  - [HIGH] CWE-79: Cross-site Scripting in src/components/Search.tsx:123

Action Required: Fix blocking findings before merge is permitted.

Other files in this skill

assets/template.md (verbatim)

SAST Pipeline Configuration Templates

GitHub Actions: Combined CodeQL + Semgrep Workflow

# .github/workflows/sast-pipeline.yml
name: "SAST Security Pipeline"

on:
  push:
    branches: [main, develop]
  pull_request:
    branches: [main]
  schedule:
    - cron: '0 3 * * 1'

concurrency:
  group: sast-${{ github.ref }}
  cancel-in-progress: true

jobs:
  # ─────────────── CodeQL Analysis ───────────────
  codeql:
    name: CodeQL (${{ matrix.language }})
    runs-on: ubuntu-latest
    permissions:
      actions: read
      contents: read
      security-events: write
    strategy:
      fail-fast: false
      matrix:
        language: ['javascript', 'python']
    steps:
      - uses: actions/checkout@v4

      - name: Initialize CodeQL
        uses: github/codeql-action/init@v3
        with:
          languages: ${{ matrix.language }}
          queries: security-extended
          config-file: .github/codeql/codeql-config.yml

      - name: Autobuild
        uses: github/codeql-action/autobuild@v3

      - name: Perform Analysis
        uses: github/codeql-action/analyze@v3
        with:
          category: "/language:${{ matrix.language }}"

  # ─────────────── Semgrep Scan ───────────────
  semgrep:
    name: Semgrep Scan
    runs-on: ubuntu-latest
    permissions:
      security-events: write
      contents: read
    container:
      image: semgrep/semgrep:latest
    steps:
      - uses: actions/checkout@v4

      - name: Run Semgrep
        run: |
          semgrep ci \
            --config auto \
            --config p/owasp-top-ten \
            --config p/cwe-top-25 \
            --config .semgrep/ \
            --sarif --output semgrep.sarif \
            --severity ERROR \
            --error
        env:
          SEMGREP_APP_TOKEN: ${{ secrets.SEMGREP_APP_TOKEN }}

      - name: Upload SARIF
        if: always()
        uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: semgrep.sarif
          category: semgrep

  # ─────────────── Quality Gate ───────────────
  security-gate:
    name: Security Quality Gate
    needs: [codeql, semgrep]
    runs-on: ubuntu-latest
    if: always()
    steps:
      - name: Check SAST Results
        run: |
          if [ "${{ needs.codeql.result }}" == "failure" ] || [ "${{ needs.semgrep.result }}" == "failure" ]; then
            echo "::error::SAST security gate failed. Review findings in the Security tab."
            exit 1
          fi
          echo "Security gate passed."

CodeQL Custom Configuration

# .github/codeql/codeql-config.yml
name: "Organization CodeQL Config"

queries:
  - uses: security-extended
  - uses: security-and-quality

paths-ignore:
  - '**/test/**'
  - '**/tests/**'
  - '**/spec/**'
  - '**/vendor/**'
  - '**/node_modules/**'
  - '**/__mocks__/**'
  - '**/*.test.js'
  - '**/*.test.ts'
  - '**/*.spec.py'
  - '**/migrations/**'

query-filters:
  - exclude:
      id: js/unused-local-variable
  - exclude:
      id: py/unused-import

Semgrep Ignore File

# .semgrepignore
# Test files
*_test.go
*_test.py
*.test.js
*.test.ts
*.spec.js
*.spec.ts
test/
tests/
__tests__/
spec/

# Generated code
*_generated.go
*.pb.go
**/generated/**

# Vendored dependencies
vendor/
node_modules/
third_party/

# Build artifacts
dist/
build/
out/

Branch Protection Configuration (Terraform)

# branch-protection.tf
resource "github_branch_protection" "main" {
  repository_id = github_repository.app.node_id
  pattern       = "main"

  required_status_checks {
    strict   = true
    contexts = [
      "CodeQL (javascript)",
      "CodeQL (python)",
      "Semgrep Scan",
      "Security Quality Gate"
    ]
  }

  required_pull_request_reviews {
    required_approving_review_count = 1
    dismiss_stale_reviews           = true
  }

  enforce_admins = true

  allows_deletions    = false
  allows_force_pushes = false
}

SARIF Report Aggregation Script

#!/bin/bash
# aggregate-sarif.sh - Merge multiple SARIF files for unified upload
set -euo pipefail

OUTPUT="merged-results.sarif"
SARIF_FILES=($(find . -name "*.sarif" -type f))

if [ ${#SARIF_FILES[@]} -eq 0 ]; then
  echo "No SARIF files found"
  exit 0
fi

# Use jq to merge SARIF runs
jq -s '{
  "$schema": "https://raw.githubusercontent.com/oasis-tcs/sarif-spec/main/sarif-2.1/schema/sarif-schema-2.1.0.json",
  "version": "2.1.0",
  "runs": [.[].runs[]]
}' "${SARIF_FILES[@]}" > "$OUTPUT"

echo "Merged ${#SARIF_FILES[@]} SARIF files into $OUTPUT"
TOTAL=$(jq '[.runs[].results | length] | add' "$OUTPUT")
echo "Total findings: $TOTAL"

references/api-reference.md (verbatim)

API Reference: SAST in GitHub Actions Pipeline

Semgrep CLI

Installation

pip install semgrep

Scan Commands

semgrep scan --config auto --json .           # Auto-detect rules
semgrep scan --config p/owasp-top-ten --json . # OWASP rules
semgrep scan --config p/ci --sarif .           # CI-optimized rules

JSON Output Structure

{"results": [{"check_id": "rule-id", "path": "file.py",
  "start": {"line": 10}, "extra": {"severity": "ERROR",
  "message": "...", "metadata": {"cwe": ["CWE-89"], "owasp": ["A03"]}}}]}

Severity Levels

Level Action
ERROR Block merge
WARNING Require review
INFO Advisory only

GitHub Actions Integration

Semgrep Action

- uses: returntocorp/semgrep-action@v1
  with:
    config: auto
    generateSarif: "1"

SARIF Upload

- uses: github/codeql-action/upload-sarif@v3
  with:
    sarif_file: semgrep.sarif

SARIF 2.1.0 Schema

Field Description
runs[].tool.driver.name Scanner name
runs[].tool.driver.rules Rule definitions
runs[].results Finding instances
results[].ruleId Matching rule ID
results[].level error, warning, note

References

references/standards.md (verbatim)

Standards Reference: SAST in GitHub Actions

OWASP SAMM - Verification: Security Testing

Maturity Level 1

  • Perform automated SAST scanning with default rulesets on all application code
  • Results are visible to development teams through IDE or CI/CD integration

Maturity Level 2

  • Customize SAST rules to reduce false positives below 20%
  • Track and triage all findings with defined SLAs per severity
  • Integrate SAST results into a centralized vulnerability management system

Maturity Level 3

  • Correlate SAST findings with DAST and SCA results for comprehensive coverage
  • Measure and improve detection accuracy through benchmarking against known vulnerabilities
  • Custom rules cover organization-specific vulnerability patterns and deprecated APIs

NIST SSDF (SP 800-218) - Produce Well-Secured Software

PW.7: Review and Analyze Code

  • PW.7.1: Determine whether SAST tools should be used and select appropriate tools
  • PW.7.2: Use SAST tools to analyze source code and identify vulnerabilities
  • Configure tools to analyze code for compliance with secure coding standards

PW.8: Test Executable Code

  • Integration of SAST into CI/CD ensures code is tested before deployment
  • Findings are tracked and remediated according to organizational policy

CIS Software Supply Chain Security Guide

Source Code (SC) Controls

  • SC-2: Enforce branch protection requiring SAST checks to pass
  • SC-3: Require code review in addition to automated scanning
  • SC-4: Automate security testing in the build pipeline

Build (BD) Controls

  • BD-1: Define and enforce security requirements for build processes
  • BD-2: Integrate multiple security testing tools for defense in depth

OWASP Top 10 Coverage Matrix

OWASP Category CodeQL Semgrep Combined
A01: Broken Access Control Partial Yes Yes
A02: Cryptographic Failures Yes Yes Yes
A03: Injection Yes Yes Yes
A04: Insecure Design No Partial Partial
A05: Security Misconfiguration Partial Yes Yes
A06: Vulnerable Components No No No (Use SCA)
A07: Auth Failures Yes Yes Yes
A08: Software Integrity No Partial Partial
A09: Logging Failures Partial Yes Yes
A10: SSRF Yes Yes Yes

PCI DSS v4.0 Mapping

  • Requirement 6.2.4: Software engineering techniques or automated methods prevent or mitigate common software attacks
  • Requirement 6.3.2: An inventory of bespoke and custom software and third-party software components is maintained
  • Requirement 6.5.4: SAST tools are run as part of the software development lifecycle

SOC 2 Trust Service Criteria

  • CC7.1: Deploy detection and monitoring mechanisms for anomalies indicative of actual or attempted attacks
  • CC8.1: Authorize, design, develop or acquire, configure, document, test, approve, and implement changes to infrastructure, data, software, and procedures

references/workflows.md (verbatim)

Workflow Reference: SAST in GitHub Actions Pipeline

End-to-End SAST Integration Workflow

Developer Push/PR
       │
       ▼
┌──────────────────┐
│ GitHub Actions    │
│ Trigger           │
└──────┬───────────┘
       │
       ├──────────────────────┐
       ▼                      ▼
┌──────────────┐    ┌──────────────┐
│ CodeQL Init  │    │ Semgrep CI   │
│ + Autobuild  │    │ + Custom     │
│ + Analyze    │    │   Rules      │
└──────┬───────┘    └──────┬───────┘
       │                    │
       ▼                    ▼
┌──────────────┐    ┌──────────────┐
│ SARIF Upload │    │ SARIF Upload │
│ (CodeQL)     │    │ (Semgrep)    │
└──────┬───────┘    └──────┬───────┘
       │                    │
       └──────────┬─────────┘
                  ▼
       ┌──────────────────┐
       │ GitHub Security  │
       │ Tab / Dashboard  │
       └──────┬───────────┘
              │
              ▼
       ┌──────────────────┐
       │ Branch Protection│
       │ Quality Gate     │
       └──────┬───────────┘
              │
    ┌─────────┴──────────┐
    ▼                    ▼
 PASS: Merge          FAIL: Block
 Permitted            + Notify Dev

CodeQL Analysis Deep Dive

Database Creation Phase

  1. CodeQL extracts source code into a relational database
  2. For compiled languages (Java, C++, C#, Go), the build process is intercepted
  3. For interpreted languages (Python, JavaScript, Ruby), source files are parsed directly
  4. The database captures the full AST, data flow, and control flow of the program

Query Execution Phase

  1. Security queries analyze the database for known vulnerability patterns
  2. Taint tracking follows data from untrusted sources to dangerous sinks
  3. Dataflow analysis tracks variable assignments across method boundaries
  4. Results are deduplicated and ranked by confidence and severity

Query Suites

  • default: Core security queries with high precision and low false positive rate
  • security-extended: Additional queries covering more vulnerability types
  • security-and-quality: All security queries plus code quality checks

Semgrep Rule Authoring Process

Rule Development Lifecycle

  1. Identify a vulnerability pattern from a recent security incident or code review
  2. Write the pattern using Semgrep syntax with pattern, pattern-either, pattern-not
  3. Test the rule against known vulnerable and safe code samples
  4. Add metadata: CWE ID, OWASP category, severity, remediation guidance
  5. Deploy via .semgrep/ directory or Semgrep App registry
  6. Monitor false positive rate and refine patterns

Pattern Operators Reference

Operator Purpose
pattern Match a single code pattern
pattern-either Match any of multiple patterns (OR)
pattern-not Exclude specific patterns from matches
pattern-inside Match only within a containing pattern
pattern-not-inside Exclude matches within a containing pattern
metavariable-regex Constrain metavariable values with regex
metavariable-comparison Compare metavariable values numerically

SARIF Processing Pipeline

SARIF Structure

{
  "$schema": "https://raw.githubusercontent.com/oasis-tcs/sarif-spec/main/sarif-2.1/schema/sarif-schema-2.1.0.json",
  "version": "2.1.0",
  "runs": [
    {
      "tool": {
        "driver": {
          "name": "Semgrep",
          "rules": []
        }
      },
      "results": [
        {
          "ruleId": "hardcoded-database-url",
          "level": "error",
          "message": { "text": "..." },
          "locations": [
            {
              "physicalLocation": {
                "artifactLocation": { "uri": "src/config.py" },
                "region": { "startLine": 42 }
              }
            }
          ]
        }
      ]
    }
  ]
}

Triage and Remediation Workflow

Severity-Based SLA

Severity Triage SLA Remediation SLA Escalation
Critical 1 business day 3 business days Security Lead + VP Eng
High 3 business days 10 business days Security Lead
Medium 5 business days 30 business days Team Lead
Low 10 business days 90 business days Backlog

Finding States

  1. Open: New finding not yet reviewed
  2. Confirmed: Finding validated as true positive
  3. False Positive: Finding dismissed with justification
  4. Fixed: Remediation committed and verified by rescan
  5. Won't Fix: Accepted risk with documented justification and risk owner

Back to mukul975/Anthropic-Cybersecurity-Skills (817 security skills) or Agent skills.