What it does. Develops precise YARA and YARA-X rules for malware detection by identifying Part of mukul975/Anthropic-Cybersecurity-Skills (817 security skills) (mukul975/Anthropic-Cybersecurity-Skills).
Install
npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill performing-yara-rule-development-for-detection, or copy the skill folder into ~/.claude/skills/performing-yara-rule-development-for-detection/.
- Raw file:
curl -sL https://raw.githubusercontent.com/mukul975/Anthropic-Cybersecurity-Skills/HEAD/skills/performing-yara-rule-development-for-detection/SKILL.md
SKILL.md (verbatim)
name: performing-yara-rule-development-for-detection
description: Develops precise YARA and YARA-X rules for malware detection by identifying
unique strings, byte sequences, PE header traits, and behavioral indicators in
unpacked malware artifacts while minimizing false positives. Use when building
detection signatures for threat hunting, classifying malware families, or authoring
rules from IOCs such as C2 URLs, mutex names, and encryption constants.
domain: cybersecurity
subdomain: malware-analysis
tags:
- yara
- malware-detection
- signature-development
- threat-hunting
- pattern-matching
- yara-x
- indicator-development
version: '1.0'
author: mahipal
license: Apache-2.0
nist_csf:
- DE.AE-02
- RS.AN-03
- ID.RA-01
- DE.CM-01
mitre_attack:
- T1027
- T1055
- T1140
- T1497
Performing YARA Rule Development for Detection
Overview
YARA is the pattern matching swiss knife for malware researchers, enabling identification and classification of malware based on textual or binary patterns. Effective YARA rules combine unique string patterns, byte sequences, PE header characteristics, import table analysis, and conditional logic to detect malware families while avoiding false positives. Modern YARA-X (rewritten in Rust, stable since June 2025) brings improved performance and new modules. Rules should target unpacked malware artifacts like hardcoded stack strings, C2 URLs, mutex names, encryption constants, and unique code sequences rather than packer signatures.
When to Use
- When conducting security assessments that involve performing yara rule development for detection
- When following incident response procedures for related security events
- When performing scheduled security testing or auditing activities
- When validating security controls through hands-on testing
Prerequisites
- Python 3.9+ with
yara-python library
- YARA 4.5+ or YARA-X 0.10+
- PE analysis tools (
pefile, pestudio)
- Hex editor for identifying unique byte patterns
- Access to malware samples (VirusTotal, MalwareBazaar)
- Understanding of PE file format, strings, and import tables
Key Concepts
Rule Structure
Every YARA rule consists of three sections: meta (optional descriptive metadata), strings (pattern definitions), and condition (matching logic). String types include text strings (ASCII/wide/nocase), hex patterns with wildcards and jumps, and regular expressions. Conditions combine string matches with file properties using boolean operators.
String Selection Strategy
Effective rules target patterns that are unique to the malware family and survive recompilation. Hardcoded stack strings are excellent choices because compilers embed them consistently. C2 domain patterns, custom encryption routines, unique error messages, and specific API call sequences provide stable detection anchors. Avoid compiler-generated boilerplate and common library strings.
YARA evaluates conditions short-circuit style. Place the most discriminating and cheapest-to-evaluate conditions first. Use filesize limits to skip irrelevant files quickly. Minimize regex usage in favor of hex patterns. Use private rules as building blocks for complex detection logic without generating standalone matches.
Workflow
Step 1: Analyze Sample for Unique Patterns
#!/usr/bin/env python3
"""Extract candidate strings and byte patterns for YARA rule creation."""
import pefile
import re
import sys
from collections import Counter
def extract_strings(filepath, min_length=6):
"""Extract ASCII and wide strings from binary."""
with open(filepath, 'rb') as f:
data = f.read()
# ASCII strings
ascii_strings = re.findall(
rb'[\x20-\x7e]{' + str(min_length).encode() + rb',}', data
)
# Wide (UTF-16LE) strings
wide_strings = re.findall(
rb'(?:[\x20-\x7e]\x00){' + str(min_length).encode() + rb',}', data
)
return {
'ascii': [s.decode('ascii') for s in ascii_strings],
'wide': [s.decode('utf-16-le') for s in wide_strings],
}
def analyze_pe_imports(filepath):
"""Extract import table for API-based detection."""
try:
pe = pefile.PE(filepath)
except pefile.PEFormatError:
return []
imports = []
if hasattr(pe, 'DIRECTORY_ENTRY_IMPORT'):
for entry in pe.DIRECTORY_ENTRY_IMPORT:
dll_name = entry.dll.decode('utf-8', errors='replace')
for imp in entry.imports:
if imp.name:
func_name = imp.name.decode('utf-8', errors='replace')
imports.append(f"{dll_name}!{func_name}")
return imports
def find_unique_byte_patterns(filepath, pattern_length=16):
"""Find unique byte sequences suitable for YARA hex patterns."""
with open(filepath, 'rb') as f:
data = f.read()
try:
pe = pefile.PE(filepath)
# Focus on code section
for section in pe.sections:
if section.Characteristics & 0x20000000: # IMAGE_SCN_MEM_EXECUTE
code_start = section.PointerToRawData
code_end = code_start + section.SizeOfRawData
code_data = data[code_start:code_end]
break
else:
code_data = data
except Exception:
code_data = data
# Find byte patterns that appear exactly once
patterns = []
for i in range(0, len(code_data) - pattern_length, 4):
pattern = code_data[i:i+pattern_length]
if pattern.count(b'\x00') < pattern_length // 3: # Skip null-heavy
hex_pattern = ' '.join(f'{b:02X}' for b in pattern)
patterns.append(hex_pattern)
# Count frequency and return unique ones
freq = Counter(patterns)
unique = [p for p, count in freq.items() if count == 1]
return unique[:20] # Top 20 candidates
def suggest_rule_strings(filepath):
"""Suggest strings and patterns for YARA rule."""
print(f"[+] Analyzing: {filepath}")
# Extract strings
strings = extract_strings(filepath)
# Filter for suspicious/unique strings
suspicious_keywords = [
'http', 'https', 'cmd', 'powershell', 'mutex', 'pipe',
'password', 'credential', 'inject', 'hook', 'debug',
'sandbox', 'virtual', 'vmware', 'vbox',
]
print("\n[+] Suspicious ASCII strings:")
for s in strings['ascii']:
if any(kw in s.lower() for kw in suspicious_keywords):
print(f" $ = \"{s}\" ascii")
print("\n[+] Suspicious wide strings:")
for s in strings['wide']:
if any(kw in s.lower() for kw in suspicious_keywords):
print(f" $ = \"{s}\" wide")
# Import analysis
imports = analyze_pe_imports(filepath)
suspicious_apis = [
'VirtualAlloc', 'VirtualProtect', 'WriteProcessMemory',
'CreateRemoteThread', 'NtUnmapViewOfSection', 'RtlMoveMemory',
'OpenProcess', 'CreateToolhelp32Snapshot',
'InternetOpenA', 'HttpSendRequestA',
'CryptEncrypt', 'CryptDecrypt',
]
print("\n[+] Suspicious imports:")
for imp in imports:
func = imp.split('!')[-1]
if func in suspicious_apis:
print(f" {imp}")
# Byte patterns
print("\n[+] Candidate hex patterns:")
patterns = find_unique_byte_patterns(filepath)
for p in patterns[:5]:
print(f" $hex = {{ {p} }}")
if __name__ == "__main__":
if len(sys.argv) < 2:
print(f"Usage: {sys.argv[0]} <sample_path>")
sys.exit(1)
suggest_rule_strings(sys.argv[1])
Step 2: Write and Test YARA Rules
import yara
import os
def create_yara_rule(rule_name, meta, strings, condition):
"""Generate a YARA rule from components."""
meta_str = "\n".join(f' {k} = "{v}"' for k, v in meta.items())
strings_str = "\n".join(f" {s}" for s in strings)
rule = f"""rule {rule_name} {{
meta:
{meta_str}
strings:
{strings_str}
condition:
{condition}
}}"""
return rule
def test_yara_rule(rule_text, test_dir):
"""Compile and test YARA rule against sample directory."""
try:
rules = yara.compile(source=rule_text)
except yara.SyntaxError as e:
print(f"[-] YARA syntax error: {e}")
return None
results = {"matches": [], "no_match": []}
for filename in os.listdir(test_dir):
filepath = os.path.join(test_dir, filename)
if not os.path.isfile(filepath):
continue
matches = rules.match(filepath)
if matches:
results["matches"].append({
"file": filename,
"rules": [m.rule for m in matches],
})
else:
results["no_match"].append(filename)
print(f"[+] Matches: {len(results['matches'])}")
print(f"[-] No match: {len(results['no_match'])}")
return results
# Example: Create a rule for a hypothetical malware family
example_rule = create_yara_rule(
rule_name="MalwareFamily_Variant_A",
meta={
"description": "Detects MalwareFamily Variant A",
"author": "Malware Analysis Team",
"date": "2025-01-01",
"hash": "abc123...",
"tlp": "WHITE",
},
strings=[
'$mutex = "Global\\\\UniqueM4lwareMutex" ascii wide',
'$c2_pattern = /https?:\\/\\/[a-z]{5,10}\\.(xyz|top|buzz)\\/gate\\.php/',
'$api1 = "VirtualAllocEx" ascii',
'$api2 = "WriteProcessMemory" ascii',
'$api3 = "CreateRemoteThread" ascii',
'$hex_decrypt = { 8B 45 ?? 33 C1 89 45 ?? 83 C1 04 }',
'$pdb = "C:\\\\Users\\\\" ascii',
],
condition=(
'uint16(0) == 0x5A4D and filesize < 2MB and '
'($mutex or $c2_pattern) and '
'2 of ($api*) and '
'$hex_decrypt'
),
)
print(example_rule)
import time
def benchmark_rule(rule_text, scan_directory, iterations=3):
"""Benchmark YARA rule scan performance."""
rules = yara.compile(source=rule_text)
files = []
for root, _, filenames in os.walk(scan_directory):
for f in filenames:
files.append(os.path.join(root, f))
print(f"[+] Benchmarking against {len(files)} files "
f"({iterations} iterations)")
times = []
for i in range(iterations):
start = time.perf_counter()
matches = 0
for filepath in files:
try:
result = rules.match(filepath)
if result:
matches += 1
except Exception:
pass
elapsed = time.perf_counter() - start
times.append(elapsed)
print(f" Iteration {i+1}: {elapsed:.3f}s ({matches} matches)")
avg_time = sum(times) / len(times)
files_per_sec = len(files) / avg_time
print(f"\n[+] Average: {avg_time:.3f}s ({files_per_sec:.0f} files/sec)")
return avg_time
Validation Criteria
- YARA rules compile without syntax errors
- Rules detect target malware family samples with zero false negatives
- False positive rate below 0.1% when scanned against clean file corpus
- Rule performance allows scanning 1000+ files per second
- Rules survive minor malware modifications (recompilation, string changes)
- Metadata includes hash, author, date, description, and TLP marking
References
Other files in this skill
assets/template.md (verbatim)
YARA Rule Development Report
| Field |
Value |
| Rule Name |
|
| Target Family |
|
| Author |
|
| Date Created |
|
| TLP |
WHITE |
Detection Targets
| Pattern Type |
Value |
Rationale |
| String |
|
|
| Hex Pattern |
|
|
| Import |
|
|
Testing Results
| Metric |
Value |
| True Positives |
|
| False Negatives |
|
| False Positives |
|
| Detection Rate |
% |
| FP Rate |
% |
Deployment Recommendations
- Deploy to endpoint scanning infrastructure
- Add to YARA retrohunt on VirusTotal
- Integrate with SIEM alerting pipeline
references/api-reference.md (verbatim)
API Reference: YARA Rule Development for Detection
yara-python API
| Method |
Description |
yara.compile(filepath=path) |
Compile rule from file |
yara.compile(source=string) |
Compile rule from string |
yara.compile(filepaths={ns: path}) |
Compile with namespaces |
rules.match(filepath=path) |
Scan file against compiled rules |
rules.match(data=bytes) |
Scan bytes in memory |
rules.match(filepath, timeout=30) |
Scan with timeout |
Match Object Attributes
| Attribute |
Description |
match.rule |
Name of matching rule |
match.namespace |
Rule namespace |
match.tags |
Rule tags list |
match.meta |
Rule metadata dict |
match.strings |
List of (offset, identifier, data) |
YARA Rule Structure
rule RuleName : tag1 tag2 {
meta:
description = "..."
author = "..."
date = "2025-01-01"
hash = "sha256_of_sample"
strings:
$s1 = "string" ascii
$s2 = "wide_string" wide
$h1 = { 4D 5A 90 00 }
$r1 = /regex[0-9]+/
condition:
uint16(0) == 0x5A4D and 3 of ($s*)
}
Condition Operators
| Operator |
Description |
X of ($s*) |
X or more strings match |
all of ($s*) |
All strings match |
any of ($s*) |
At least one matches |
uint16(0) == 0x5A4D |
PE file magic bytes |
filesize < 10MB |
File size constraint |
Python Libraries
| Library |
Version |
Purpose |
yara-python |
>=4.3 |
Compile and scan YARA rules |
hashlib |
stdlib |
SHA256 of samples |
re |
stdlib |
String extraction |
References
references/standards.md (verbatim)
YARA Rule Development Standards
Rule Naming Convention
Malware_Family_Variant: For specific malware variants
APT_Group_Tool: For threat actor associated tools
Exploit_CVE_YYYY_NNNN: For exploit payloads
Technique_Name: For generic technique detection
Rule Quality Metrics
| Metric |
Target |
Description |
| True Positive Rate |
>99% |
Detection of known samples |
| False Positive Rate |
<0.1% |
Matches on clean files |
| Scan Speed |
>1000 files/s |
Processing performance |
| Maintenance Burden |
Low |
Frequency of updates needed |
String Types Reference
| Type |
Syntax |
Use Case |
| ASCII text |
"text" ascii |
Plain text strings |
| Wide text |
"text" wide |
UTF-16LE encoded strings |
| Case-insensitive |
"text" nocase |
Variable casing |
| Hex pattern |
{ AA BB CC } |
Byte sequences |
| Wildcard hex |
{ AA ?? CC } |
Single byte wildcard |
| Jump hex |
{ AA [2-4] CC } |
Variable length gap |
| Regex |
/pattern/ |
Complex pattern matching |
MITRE ATT&CK Relevance
- T1027 - Obfuscated Files: Rules detect packed/encoded malware
- T1036 - Masquerading: Rules identify file mimicry
- T1059 - Command Interpreter: Rules detect malicious scripts
References
references/workflows.md (verbatim)
YARA Rule Development Workflows
Workflow 1: Sample-Driven Rule Creation
[Malware Sample] --> [Static Analysis] --> [Extract Unique Strings] --> [Draft Rule]
|
v
[Test Against Samples]
|
v
[Test Against Clean Files]
|
v
[Deploy to Production]
Workflow 2: Family-Wide Detection
[Multiple Samples] --> [Cross-Sample Analysis] --> [Find Common Patterns]
|
v
[Build Generic Rule]
|
v
[Validate Coverage]
Workflow 3: Threat Hunt Integration
[Intelligence Report] --> [Extract IOCs] --> [Convert to YARA] --> [Retrohunt]
|
v
[Triage New Matches]
Back to mukul975/Anthropic-Cybersecurity-Skills (817 security skills) or Agent skills.