Files
giglez/scripts/benchmark_real.py
T
leetcrypt 6945ece853 test: add real-world Flipper Zero capture benchmark - CRITICAL FINDINGS
## Summary

Added real-world testing infrastructure with 11 actual Flipper Zero community captures. Results reveal **CRITICAL GAP**: 0% accuracy on real captures vs 33% on synthetic signals.

## What Was Added

### Test Infrastructure
- **tests/real_captures/** - 11 real .sub files from UberGuidoZ/Flipper repository
  - 3 weather sensors (LaCrosse, Acurite, Nexus)
  - 2 garage doors (LiftMaster 433MHz, Security+ 2.0)
  - 2 doorbells (GE, Byron)
  - 1 LED remote
  - 3 ceiling fan controls

- **tests/real_captures/manifest.json** - Test case metadata with expected protocols

- **scripts/benchmark_real.py** - Real-world benchmark runner
  - Loads manifest.json
  - Runs identification on each capture
  - Generates REAL_TEST_RESULTS.md
  - Compares with synthetic results

### Documentation
- **REAL_TEST_RESULTS.md** - Benchmark results (0% accuracy)
- **docs/REAL_CAPTURE_ANALYSIS.md** - Comprehensive gap analysis

## Results

### Performance Comparison

| Metric | Synthetic | Real-World | Delta |
|--------|-----------|------------|-------|
| Top-1 Accuracy | 33.3% | **0.0%** | **-33.3%** |
| Top-3 Accuracy | 50.0% | **0.0%** | **-50.0%** |
| High Confidence | 66.7% | 9.1% | -57.6% |
| No Matches | 8.3% | **45.5%** | +37.2% |

### Root Causes Identified

1. **Decoded File Format (45% of failures)** - CRITICAL
   - 5/11 files are already decoded (Protocol: Holtek_HT12X), not RAW
   - Our system ONLY processes RAW_Data
   - All ceiling fan and LED remote files fail immediately
   - **Fix**: Add support for parsed decode .sub file format

2. **Missing Protocols (27% of failures)**
   - Acurite 02077M not in database (we have 609TXC)
   - GE Doorbell 19297 not in database
   - Byron DB421E not in database
   - LiftMaster not in database
   - **Fix**: Add these protocols from real captures

3. **Real Signal Complexity (27% of failures)**
   - LaCrosse real capture: 131 pulses (multi-packet)
   - Synthetic LaCrosse: 40 bits (single packet)
   - Real signals have noise, jitter, interference
   - **Fix**: Packet segmentation, higher noise tolerance

4. **Protocol Family Competition (9% of failures)**
   - Nexus found at rank 34 (beaten by Oregon Scientific)
   - Similar protocols competing instead of grouping
   - **Fix**: Protocol family scoring

## Key Insights

### Why Synthetic Worked (33% accuracy):
- Perfect timing with controlled 5-15% jitter
- Single packet per file
- All RAW format
- Known protocol parameters

### Why Real Failed (0% accuracy):
- 45% already decoded (not RAW)
- 27% missing from database
- Variable signal quality
- Multi-packet transmissions
- Real-world interference

**Conclusion**: System was optimized for unrealistic synthetic signals.

## Recommendations (Priority Order)

1. **P1 - Support Decoded Files** (+45% accuracy)
   - Parse Protocol, Bit, Key, TE fields
   - Match by protocol name + timing element
   - Skip RAW analysis for decoded files

2. **P2 - Add Missing Protocols** (+27% accuracy)
   - Import Acurite 02077M, GE/Byron doorbells, LiftMaster, Holtek HT12X
   - Source from real captures or RTL_433 updates

3. **P3 - Improve Robustness** (+18% accuracy)
   - Packet segmentation for multi-transmission captures
   - Increase jitter tolerance to 25%
   - Repetition detection

4. **P4 - Family Grouping** (+9% accuracy)
   - Group similar protocols (Nexus/Oregon/Acurite)
   - Boost exact name matches

**Expected Final Accuracy**: 60-70% after P1+P2 implementation.

## Files Changed

- tests/real_captures/*.sub (11 files, 192KB)
- tests/real_captures/manifest.json (11 test cases)
- scripts/benchmark_real.py (370 lines)
- REAL_TEST_RESULTS.md (benchmark results)
- docs/REAL_CAPTURE_ANALYSIS.md (comprehensive analysis)

🎯 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-02-15 18:08:48 -08:00

411 lines
15 KiB
Python
Executable File

#!/usr/bin/env python3
"""
Real-World RF Device Identification Benchmark
Tests the identification system against actual Flipper Zero community captures.
Compares performance on real data vs synthetic data.
"""
import sys
import json
from pathlib import Path
from typing import List, Dict, Tuple
from dataclasses import dataclass, field
from collections import defaultdict
# Add project root to path
sys.path.insert(0, str(Path(__file__).parent.parent))
from src.parser.sub_parser import parse_sub_file
from src.matcher.pattern_decoder import get_pattern_decoder, DeviceMatch
@dataclass
class RealCaptureResult:
"""Result for a single real capture test"""
filename: str
expected_protocol: str
category: str
frequency: int
# Results
top_match: str = ""
top_confidence: float = 0.0
all_matches: List[DeviceMatch] = field(default_factory=list)
# Metrics
correct: bool = False
rank: int = -1 # Rank of expected protocol (-1 if not found)
# Performance
parse_time_ms: float = 0.0
match_time_ms: float = 0.0
@dataclass
class RealBenchmarkSummary:
"""Overall benchmark statistics for real captures"""
total_tests: int = 0
# Top-K accuracy
top1_accuracy: float = 0.0
top3_accuracy: float = 0.0
top5_accuracy: float = 0.0
# Confidence distribution
high_confidence_count: int = 0 # >80%
medium_confidence_count: int = 0 # 50-80%
low_confidence_count: int = 0 # <50%
none_found_count: int = 0 # No matches
# Category breakdown
category_stats: Dict[str, Dict] = field(default_factory=dict)
# Performance
avg_parse_time_ms: float = 0.0
avg_match_time_ms: float = 0.0
# All results
results: List[RealCaptureResult] = field(default_factory=list)
class RealWorldBenchmark:
"""Benchmark runner for real-world captures"""
def __init__(self):
self.decoder = get_pattern_decoder()
def run_benchmark(
self,
test_cases_dir: Path,
manifest_path: Path
) -> RealBenchmarkSummary:
"""
Run benchmark on real-world captures
Args:
test_cases_dir: Directory containing .sub files
manifest_path: Path to manifest.json
Returns:
RealBenchmarkSummary with all metrics
"""
# Load manifest
with open(manifest_path) as f:
manifest = json.load(f)
test_cases = manifest['test_cases']
summary = RealBenchmarkSummary()
summary.total_tests = len(test_cases)
print(f"Running real-world benchmark on {len(test_cases)} captures...")
print()
for i, test_case in enumerate(test_cases, 1):
filepath = test_cases_dir / test_case['filename']
print(f"[{i}/{len(test_cases)}] Testing {filepath.name}...", end=" ")
result = self._test_single_file(
filepath,
test_case['expected_protocol'],
test_case['category'],
test_case['frequency']
)
summary.results.append(result)
# Update category stats
if result.category not in summary.category_stats:
summary.category_stats[result.category] = {
'total': 0,
'correct': 0,
'top3': 0,
'top5': 0
}
stats = summary.category_stats[result.category]
stats['total'] += 1
if result.correct and result.rank == 1:
stats['correct'] += 1
print(f"✓ CORRECT (confidence: {result.top_confidence:.1%})")
elif result.rank > 0 and result.rank <= 3:
stats['top3'] += 1
print(f"⚠ Rank {result.rank} (expected in top 3)")
elif result.rank > 0 and result.rank <= 5:
stats['top5'] += 1
print(f"⚠ Rank {result.rank} (expected in top 5)")
else:
if result.top_match:
print(f"✗ WRONG: {result.top_match} (confidence: {result.top_confidence:.1%})")
else:
print(f"✗ NO MATCHES FOUND")
print()
# Calculate aggregate metrics
self._calculate_summary_metrics(summary)
return summary
def _test_single_file(
self,
filepath: Path,
expected_protocol: str,
category: str,
frequency: int
) -> RealCaptureResult:
"""Test a single .sub file"""
import time
result = RealCaptureResult(
filename=filepath.name,
expected_protocol=expected_protocol,
category=category,
frequency=frequency
)
# Parse file
t0 = time.time()
try:
signal_metadata = parse_sub_file(str(filepath))
except Exception as e:
print(f"Parse error: {e}")
return result
result.parse_time_ms = (time.time() - t0) * 1000
# Run matching
t0 = time.time()
try:
matches = self.decoder.decode(signal_metadata)
result.all_matches = matches
except Exception as e:
print(f"Match error: {e}")
return result
result.match_time_ms = (time.time() - t0) * 1000
if not matches:
return result
# Extract top match
result.top_match = matches[0].name
result.top_confidence = matches[0].confidence
# Check if expected protocol is in results (fuzzy matching)
for rank, match in enumerate(matches, 1):
if self._protocol_matches(match.name, expected_protocol):
result.correct = (rank == 1)
result.rank = rank
break
return result
def _protocol_matches(self, actual: str, expected: str) -> bool:
"""Check if protocol names match (fuzzy matching)"""
actual_lower = actual.lower()
expected_lower = expected.lower()
# Exact match
if actual_lower == expected_lower:
return True
# Partial match
actual_parts = set(actual_lower.replace('-', ' ').split())
expected_parts = set(expected_lower.replace('-', ' ').split())
# If 2+ common tokens, consider match
common = actual_parts & expected_parts
if len(common) >= 2:
return True
# Check if expected is substring of actual
if expected_lower in actual_lower:
return True
return False
def _calculate_summary_metrics(self, summary: RealBenchmarkSummary):
"""Calculate aggregate metrics from individual results"""
if summary.total_tests == 0:
return
# Top-K accuracy
top1_correct = sum(1 for r in summary.results if r.correct and r.rank == 1)
top3_correct = sum(1 for r in summary.results if r.rank > 0 and r.rank <= 3)
top5_correct = sum(1 for r in summary.results if r.rank > 0 and r.rank <= 5)
summary.top1_accuracy = top1_correct / summary.total_tests
summary.top3_accuracy = top3_correct / summary.total_tests
summary.top5_accuracy = top5_correct / summary.total_tests
# Confidence distribution
for result in summary.results:
if result.top_confidence == 0.0:
summary.none_found_count += 1
elif result.top_confidence >= 0.8:
summary.high_confidence_count += 1
elif result.top_confidence >= 0.5:
summary.medium_confidence_count += 1
else:
summary.low_confidence_count += 1
# Performance
if summary.results:
summary.avg_parse_time_ms = sum(r.parse_time_ms for r in summary.results) / len(summary.results)
summary.avg_match_time_ms = sum(r.match_time_ms for r in summary.results) / len(summary.results)
def print_summary(self, summary: RealBenchmarkSummary):
"""Print human-readable summary"""
print("=" * 80)
print("REAL-WORLD BENCHMARK SUMMARY")
print("=" * 80)
print()
print(f"Total Tests: {summary.total_tests}")
print()
print("TOP-K ACCURACY:")
print(f" Top-1: {summary.top1_accuracy:.1%} ({int(summary.top1_accuracy * summary.total_tests)}/{summary.total_tests})")
print(f" Top-3: {summary.top3_accuracy:.1%} ({int(summary.top3_accuracy * summary.total_tests)}/{summary.total_tests})")
print(f" Top-5: {summary.top5_accuracy:.1%} ({int(summary.top5_accuracy * summary.total_tests)}/{summary.total_tests})")
print()
print("CONFIDENCE DISTRIBUTION:")
print(f" High (>80%): {summary.high_confidence_count} ({summary.high_confidence_count/summary.total_tests:.1%})")
print(f" Medium (50-80%): {summary.medium_confidence_count} ({summary.medium_confidence_count/summary.total_tests:.1%})")
print(f" Low (<50%): {summary.low_confidence_count} ({summary.low_confidence_count/summary.total_tests:.1%})")
print(f" None found: {summary.none_found_count} ({summary.none_found_count/summary.total_tests:.1%})")
print()
print("PERFORMANCE:")
print(f" Avg Parse Time: {summary.avg_parse_time_ms:.2f} ms")
print(f" Avg Match Time: {summary.avg_match_time_ms:.2f} ms")
print(f" Total Time: {summary.avg_parse_time_ms + summary.avg_match_time_ms:.2f} ms")
print()
print("BY CATEGORY:")
print(f"{'Category':<20} {'Tests':>6} {'Accuracy':>10} {'Top-3':>10}")
print("-" * 80)
for category, stats in sorted(summary.category_stats.items()):
accuracy = stats['correct'] / stats['total'] if stats['total'] > 0 else 0
top3 = stats['top3'] / stats['total'] if stats['total'] > 0 else 0
print(f"{category:<20} {stats['total']:>6} {accuracy:>9.1%} {top3:>9.1%}")
print()
def save_report(self, summary: RealBenchmarkSummary, output_path: Path):
"""Save detailed report to markdown"""
with open(output_path, 'w') as f:
f.write("# RF Device Identification - Real-World Benchmark Results\n\n")
f.write(f"**Date**: {Path(__file__).stat().st_mtime}\n")
f.write(f"**Total Tests**: {summary.total_tests}\n")
f.write(f"**Source**: Real Flipper Zero community captures\n\n")
f.write("## Overall Metrics\n\n")
f.write("### Top-K Accuracy\n\n")
f.write(f"- **Top-1**: {summary.top1_accuracy:.1%}\n")
f.write(f"- **Top-3**: {summary.top3_accuracy:.1%}\n")
f.write(f"- **Top-5**: {summary.top5_accuracy:.1%}\n\n")
f.write("### Confidence Distribution\n\n")
f.write(f"- **High (>80%)**: {summary.high_confidence_count} ({summary.high_confidence_count/summary.total_tests:.1%})\n")
f.write(f"- **Medium (50-80%)**: {summary.medium_confidence_count} ({summary.medium_confidence_count/summary.total_tests:.1%})\n")
f.write(f"- **Low (<50%)**: {summary.low_confidence_count} ({summary.low_confidence_count/summary.total_tests:.1%})\n")
f.write(f"- **None found**: {summary.none_found_count} ({summary.none_found_count/summary.total_tests:.1%})\n\n")
f.write("### Performance\n\n")
f.write(f"- **Avg Parse Time**: {summary.avg_parse_time_ms:.2f} ms\n")
f.write(f"- **Avg Match Time**: {summary.avg_match_time_ms:.2f} ms\n")
f.write(f"- **Total**: {summary.avg_parse_time_ms + summary.avg_match_time_ms:.2f} ms\n\n")
f.write("## By Category\n\n")
f.write("| Category | Tests | Top-1 Acc | Top-3 Acc |\n")
f.write("|----------|-------|-----------|------------|\n")
for category, stats in sorted(summary.category_stats.items()):
accuracy = stats['correct'] / stats['total'] if stats['total'] > 0 else 0
top3 = stats['top3'] / stats['total'] if stats['total'] > 0 else 0
f.write(f"| {category} | {stats['total']} | {accuracy:.1%} | {top3:.1%} |\n")
f.write("\n## Detailed Results\n\n")
for result in summary.results:
status = "✓" if result.correct else "✗"
f.write(f"### {status} {result.filename}\n\n")
f.write(f"- **Expected**: {result.expected_protocol}\n")
f.write(f"- **Category**: {result.category}\n")
f.write(f"- **Got**: {result.top_match if result.top_match else 'No matches'} ")
f.write(f"(confidence: {result.top_confidence:.1%})\n")
f.write(f"- **Rank**: {result.rank if result.rank > 0 else 'Not Found'}\n")
if result.all_matches:
f.write(f"\n**Top 5 Matches**:\n")
for i, match in enumerate(result.all_matches[:5], 1):
f.write(f"{i}. {match.name} ({match.confidence:.1%})\n")
f.write("\n")
def main():
"""Main benchmark runner"""
print("=" * 80)
print("REAL-WORLD RF DEVICE IDENTIFICATION BENCHMARK")
print("=" * 80)
print()
# Setup paths
test_dir = Path(__file__).parent.parent / "tests" / "real_captures"
manifest_path = test_dir / "manifest.json"
if not manifest_path.exists():
print(f"❌ Manifest not found: {manifest_path}")
return 1
# Run benchmark
benchmark = RealWorldBenchmark()
summary = benchmark.run_benchmark(test_dir, manifest_path)
# Print results
print()
benchmark.print_summary(summary)
# Save report
report_path = Path(__file__).parent.parent / "REAL_TEST_RESULTS.md"
benchmark.save_report(summary, report_path)
print(f"Detailed report saved to: {report_path}")
print()
# Compare with synthetic
print("=" * 80)
print("COMPARISON: Real vs Synthetic")
print("=" * 80)
print()
synthetic_results_path = Path(__file__).parent.parent / "TEST_RESULTS_SUMMARY.md"
if synthetic_results_path.exists():
print("Real-world results compared with synthetic benchmark:")
print(f" Real captures: {summary.total_tests} tests")
print(f" Synthetic signals: 12 tests (from TEST_RESULTS_SUMMARY.md)")
print()
print(f" Real Top-1: {summary.top1_accuracy:.1%}")
print(f" Synthetic Top-1: 33.3% (from previous run)")
print()
print(f" Real Top-3: {summary.top3_accuracy:.1%}")
print(f" Synthetic Top-3: 50.0%")
print()
# Return exit code
if summary.top1_accuracy >= 0.25:
print("✓ Benchmark PASSED (Top-1 accuracy >= 25%)")
return 0
else:
print(f"✗ Benchmark FAILED (Top-1 accuracy {summary.top1_accuracy:.1%} < 25%)")
return 1
if __name__ == '__main__':
sys.exit(main())