detecting-insider-data-exfiltration-via-dlp skill (Anthropic-Cybersecurity-Skills)
From Public Agent Wiki
Contents
What it does. 'Detects insider data exfiltration by analyzing DLP policy violations, Part of mukul975/Anthropic-Cybersecurity-Skills (817 security skills) (mukul975/Anthropic-Cybersecurity-Skills).
| Upstream | mukul975/Anthropic-Cybersecurity-Skills |
| Skill file | skills/detecting-insider-data-exfiltration-via-dlp/SKILL.md |
| License | Apache-2.0 (skill folder LICENSE) |
| Author | mukul975 |
| Fetched | 2026-09-10 |
Install
npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-insider-data-exfiltration-via-dlp, or copy the skill folder into~/.claude/skills/detecting-insider-data-exfiltration-via-dlp/.- Raw file:
curl -sL https://raw.githubusercontent.com/mukul975/Anthropic-Cybersecurity-Skills/HEAD/skills/detecting-insider-data-exfiltration-via-dlp/SKILL.md
SKILL.md (verbatim)
name: detecting-insider-data-exfiltration-via-dlp
description: 'Detects insider data exfiltration by analyzing DLP policy violations,
file access patterns, upload volume anomalies, and off-hours activity in endpoint
and cloud logs. Uses pandas for behavioral analytics and statistical baselines.
Use when investigating insider threats or building user behavior analytics for data
loss prevention.
'
domain: cybersecurity
subdomain: security-operations
tags:
- insider-threat
- data-loss-prevention
- dlp
- exfiltration-detection
- ueba
- security-operations
version: '1.0'
author: mahipal
license: Apache-2.0
nist_csf:
- DE.CM-01
- RS.MA-01
- GV.OV-01
- DE.AE-02
mitre_attack:
- T1078
- T1190
- T1059
- T1048
- T1041
Detecting Insider Data Exfiltration via DLP
When to Use
- When investigating security incidents that require detecting insider data exfiltration via dlp
- When building detection rules or threat hunting queries for this domain
- When SOC analysts need structured procedures for this analysis type
- When validating security monitoring coverage for related attack techniques
Prerequisites
- Familiarity with security operations concepts and tools
- Access to a test or lab environment for safe execution
- Python 3.8+ with required dependencies installed
- Appropriate authorization for any testing activities
Instructions
Analyze endpoint activity logs, cloud storage access, and email DLP events to detect data exfiltration patterns using behavioral baselines and statistical anomaly detection.
import pandas as pd
df = pd.read_csv("file_activity.csv", parse_dates=["timestamp"])
# Baseline: average daily upload volume per user
baseline = df.groupby(["user", df["timestamp"].dt.date])["bytes_transferred"].sum()
user_avg = baseline.groupby("user").mean()
# Alert on users exceeding 3x their baseline
today = df[df["timestamp"].dt.date == pd.Timestamp.today().date()]
today_totals = today.groupby("user")["bytes_transferred"].sum()
anomalies = today_totals[today_totals > user_avg * 3]
Key indicators:
- Upload volume exceeding 3x daily baseline
- Access to files outside normal scope
- Bulk downloads before resignation
- Off-hours file access patterns
- USB/external device usage spikes
Examples
# Detect off-hours activity
df["hour"] = df["timestamp"].dt.hour
off_hours = df[(df["hour"] < 6) | (df["hour"] > 22)]
suspicious = off_hours.groupby("user").size().sort_values(ascending=False)
Other files in this skill
references/api-reference.md (verbatim)
API Reference: Detecting Insider Data Exfiltration via DLP
Pandas Behavioral Analytics
import pandas as pd
df = pd.read_csv("activity.csv", parse_dates=["timestamp"])
# Columns: timestamp, user, action, file_path, bytes_transferred, destination
# Daily volume baseline per user
daily = df.groupby(["user", df["timestamp"].dt.date])["bytes_transferred"].sum()
baseline = daily.groupby("user").agg(["mean", "std"])
# Off-hours detection
df["hour"] = df["timestamp"].dt.hour
off_hours = df[(df["hour"] < 6) | (df["hour"] >= 22)]
# Bulk download detection
df.set_index("timestamp").groupby("user").resample("1h").size()
Exfiltration Indicators
| Indicator | Threshold | Severity |
|---|---|---|
| Volume > 3x baseline | Per user daily avg | HIGH |
| Volume > 5x baseline | Per user daily avg | CRITICAL |
| Off-hours events | > 10 per user | HIGH |
| Bulk downloads | > 50 files/hour | CRITICAL |
| USB transfers | Any volume | HIGH |
| Sensitive file access | Pattern match | HIGH |
Sensitive File Patterns
patterns = [
r"\.pem$", r"\.key$", r"\.env$",
r"credentials", r"password", r"\.kdbx$",
r"financial", r"payroll", r"customer.*data"
]
Microsoft Purview DLP API
import requests
headers = {"Authorization": "Bearer <token>"}
resp = requests.get(
"https://graph.microsoft.com/v1.0/security/alerts_v2",
headers=headers,
params={"$filter": "category eq 'DataLossPrevention'"}
)
References
- Microsoft Purview DLP: https://learn.microsoft.com/en-us/purview/dlp-learn-about-dlp
- pandas: https://pandas.pydata.org/docs/
- UEBA: https://www.gartner.com/en/information-technology/glossary/user-entity-behavior-analytics
Back to mukul975/Anthropic-Cybersecurity-Skills (817 security skills) or Agent skills.