detecting-insider-data-exfiltration-via-dlp skill (Anthropic-Cybersecurity-Skills)

From Public Agent Wiki

What it does. 'Detects insider data exfiltration by analyzing DLP policy violations, Part of mukul975/Anthropic-Cybersecurity-Skills (817 security skills) (mukul975/Anthropic-Cybersecurity-Skills).

Upstream mukul975/Anthropic-Cybersecurity-Skills
Skill file skills/detecting-insider-data-exfiltration-via-dlp/SKILL.md
License Apache-2.0 (skill folder LICENSE)
Author mukul975
Fetched 2026-09-10

Install

  • npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-insider-data-exfiltration-via-dlp, or copy the skill folder into ~/.claude/skills/detecting-insider-data-exfiltration-via-dlp/.
  • Raw file: curl -sL https://raw.githubusercontent.com/mukul975/Anthropic-Cybersecurity-Skills/HEAD/skills/detecting-insider-data-exfiltration-via-dlp/SKILL.md

SKILL.md (verbatim)

name: detecting-insider-data-exfiltration-via-dlp
description: 'Detects insider data exfiltration by analyzing DLP policy violations,
  file access patterns, upload volume anomalies, and off-hours activity in endpoint
  and cloud logs. Uses pandas for behavioral analytics and statistical baselines.
  Use when investigating insider threats or building user behavior analytics for data
  loss prevention.

  '
domain: cybersecurity
subdomain: security-operations
tags:
- insider-threat
- data-loss-prevention
- dlp
- exfiltration-detection
- ueba
- security-operations
version: '1.0'
author: mahipal
license: Apache-2.0
nist_csf:
- DE.CM-01
- RS.MA-01
- GV.OV-01
- DE.AE-02
mitre_attack:
- T1078
- T1190
- T1059
- T1048
- T1041

Detecting Insider Data Exfiltration via DLP

When to Use

  • When investigating security incidents that require detecting insider data exfiltration via dlp
  • When building detection rules or threat hunting queries for this domain
  • When SOC analysts need structured procedures for this analysis type
  • When validating security monitoring coverage for related attack techniques

Prerequisites

  • Familiarity with security operations concepts and tools
  • Access to a test or lab environment for safe execution
  • Python 3.8+ with required dependencies installed
  • Appropriate authorization for any testing activities

Instructions

Analyze endpoint activity logs, cloud storage access, and email DLP events to detect data exfiltration patterns using behavioral baselines and statistical anomaly detection.

import pandas as pd

df = pd.read_csv("file_activity.csv", parse_dates=["timestamp"])
# Baseline: average daily upload volume per user
baseline = df.groupby(["user", df["timestamp"].dt.date])["bytes_transferred"].sum()
user_avg = baseline.groupby("user").mean()

# Alert on users exceeding 3x their baseline
today = df[df["timestamp"].dt.date == pd.Timestamp.today().date()]
today_totals = today.groupby("user")["bytes_transferred"].sum()
anomalies = today_totals[today_totals > user_avg * 3]

Key indicators:

  1. Upload volume exceeding 3x daily baseline
  2. Access to files outside normal scope
  3. Bulk downloads before resignation
  4. Off-hours file access patterns
  5. USB/external device usage spikes

Examples

# Detect off-hours activity
df["hour"] = df["timestamp"].dt.hour
off_hours = df[(df["hour"] < 6) | (df["hour"] > 22)]
suspicious = off_hours.groupby("user").size().sort_values(ascending=False)

Other files in this skill

references/api-reference.md (verbatim)

API Reference: Detecting Insider Data Exfiltration via DLP

Pandas Behavioral Analytics

import pandas as pd

df = pd.read_csv("activity.csv", parse_dates=["timestamp"])
# Columns: timestamp, user, action, file_path, bytes_transferred, destination

# Daily volume baseline per user
daily = df.groupby(["user", df["timestamp"].dt.date])["bytes_transferred"].sum()
baseline = daily.groupby("user").agg(["mean", "std"])

# Off-hours detection
df["hour"] = df["timestamp"].dt.hour
off_hours = df[(df["hour"] < 6) | (df["hour"] >= 22)]

# Bulk download detection
df.set_index("timestamp").groupby("user").resample("1h").size()

Exfiltration Indicators

Indicator Threshold Severity
Volume > 3x baseline Per user daily avg HIGH
Volume > 5x baseline Per user daily avg CRITICAL
Off-hours events > 10 per user HIGH
Bulk downloads > 50 files/hour CRITICAL
USB transfers Any volume HIGH
Sensitive file access Pattern match HIGH

Sensitive File Patterns

patterns = [
    r"\.pem$", r"\.key$", r"\.env$",
    r"credentials", r"password", r"\.kdbx$",
    r"financial", r"payroll", r"customer.*data"
]

Microsoft Purview DLP API

import requests
headers = {"Authorization": "Bearer <token>"}
resp = requests.get(
    "https://graph.microsoft.com/v1.0/security/alerts_v2",
    headers=headers,
    params={"$filter": "category eq 'DataLossPrevention'"}
)

References

Back to mukul975/Anthropic-Cybersecurity-Skills (817 security skills) or Agent skills.