File size: 5,520 Bytes
e2aeac7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
---
tags:
  - robotics
  - anima
  - def-roboticattack
  - adversarial-defense
  - vla
  - robot-flow-labs
  - patch-detection
library_name: pytorch
pipeline_tag: image-classification
license: apache-2.0
datasets:
  - lerobot/libero
  - HuggingFaceVLA/smol-libero
model-index:
  - name: DEF-roboticattack-v1
    results:
      - task:
          type: image-classification
          name: Adversarial Patch Detection
        metrics:
          - type: accuracy
            value: 0.986
          - type: f1
            value: 0.986
          - type: precision
            value: 0.987
          - type: recall
            value: 0.985
---

# DEF-roboticattack β€” Adversarial Patch Detector for VLA Robotic Systems

Part of the [ANIMA Perception Suite](https://github.com/RobotFlow-Labs) by Robot Flow Labs.

## Paper

**On the Adversarial Vulnerability of Vision-Language-Action Models for Robotic Manipulation** (ICCV 2025)

William Wang et al. β€” arXiv:2411.13587

This module provides the **defense** counterpart to the paper's UADA/UPA/TMA adversarial attacks on VLA models like OpenVLA.

## Architecture

**PatchDetectorNet** β€” Lightweight multi-branch CNN (19,841 parameters) that detects adversarial patches in VLA image inputs through three parallel analysis branches:

| Branch | Purpose | Filters | Output |
|--------|---------|---------|--------|
| **Frequency** | High-pass anomaly detection | 16 | 16-d |
| **Edge** | Multi-scale edge energy (3Γ—3 + 5Γ—5) | 8+8 | 16-d |
| **Spatial** | Patch boundary consistency (7β†’3β†’3) | 16β†’32β†’32 | 32-d |
| **Classifier** | Binary patch detection | 64β†’32β†’1 | logit |

Input: `[B, 3, 224, 224]` float32 β€” Output: `[B, 1]` (sigmoid β†’ patch probability)

## Results

Trained on 85,000 real images (80K LIBERO robot manipulation frames + 5K COCO val2017) with paper-accurate UADA/UPA/TMA adversarial patches including geometric transforms (rotation + shear).

| Metric | Value |
|--------|-------|
| **Accuracy** | 98.6% |
| **Precision** | 98.7% |
| **Recall** | 98.5% |
| **F1 Score** | 0.986 |
| TP / FP / TN / FN | 2939 / 38 / 2979 / 44 |

### Inference Performance (NVIDIA L4)

| Metric | Model Only | Full Pipeline |
|--------|-----------|---------------|
| Latency (mean) | 1.50 ms | 3.13 ms |
| Throughput | 5,351 samples/s | 2,557 samples/s |

### CUDA Kernels

Two custom CUDA kernels compiled for sm_89 (L4/Ada):
- `fused_patch_apply`: 25.5Γ— speedup over CPU for adversarial patch application
- `fused_action_perturb`: PGD action-space perturbation with sign-correct zero-gradient handling

## Exported Formats

| Format | File | Size | Use Case |
|--------|------|------|----------|
| PyTorch (.pth) | `pytorch/def_roboticattack_v1.pth` | 91.7 KB | Training, fine-tuning |
| SafeTensors | `pytorch/def_roboticattack_v1.safetensors` | 81.7 KB | Fast loading, safe deployment |
| ONNX | `onnx/def_roboticattack_v1.onnx` | 81.4 KB | Cross-platform inference |
| TensorRT FP16 | `tensorrt/def_roboticattack_v1_fp16.trt` | 264 KB | Edge deployment (Jetson/L4) |
| TensorRT FP32 | `tensorrt/def_roboticattack_v1_fp32.trt` | 305 KB | Full precision inference |

## Usage

### PyTorch
```python
import torch
from safetensors.torch import load_file

# Load model
state = load_file("pytorch/def_roboticattack_v1.safetensors")
# See src/def_roboticattack/models/patch_detector.py for PatchDetectorNet class
model.load_state_dict(state)
model.eval()

# Inference
images = torch.rand(1, 3, 224, 224).cuda()
prob = torch.sigmoid(model(images))  # probability of adversarial patch
```

### Full Defense Pipeline
```python
from def_roboticattack.pipeline.runtime import DefenseRuntime

runtime = DefenseRuntime(backend="cuda")
runtime.load_model("pytorch/def_roboticattack_v1.pth")

result = runtime.full_defense(images)
# result["combined_risk"] β€” float in [0, 1]
# result["neural"]["flagged"] β€” list of booleans per image
```

### ONNX Runtime
```python
import onnxruntime as ort
import numpy as np

session = ort.InferenceSession("onnx/def_roboticattack_v1.onnx")
image = np.random.randn(1, 3, 224, 224).astype(np.float32)
logit = session.run(None, {"image": image})[0]
```

## Training

- **Hardware**: NVIDIA L4 (23 GB VRAM)
- **Framework**: PyTorch 2.11 + CUDA 12.8
- **Optimizer**: AdamW (lr=5e-4, weight_decay=0.01)
- **Scheduler**: Warmup cosine (500 warmup steps)
- **Precision**: FP16 mixed precision
- **Batch size**: 512 (auto-detected)
- **Epochs**: 30 (3,420 seconds)
- **Config**: See `configs/training.toml`

### Attack Types Used for Training

| Attack | Description | Paper Section |
|--------|-------------|---------------|
| **UADA** | Universal adversarial patch with geometric jitter | Β§3.2 |
| **UPA** | Untargeted high-frequency checkerboard patch | Β§3.3 |
| **TMA** | Targeted manipulation with directional gradients | Β§3.4 |

All patches applied with random position, rotation (Β±30Β°), shear (Β±0.2), and alpha blending (0.7–1.0).

## Defense Pipeline

The full defense stack includes:
1. **Input sanitization**: intensity clamping + Gaussian blur to suppress patch edges
2. **Heuristic detection**: Sobel edge energy with fixed threshold
3. **Neural detection**: PatchDetectorNet binary classifier
4. **Risk aggregation**: max(heuristic, neural) β†’ combined risk score [0, 1]

## Repository

- **GitHub**: [RobotFlow-Labs/DEF-roboticattack](https://github.com/RobotFlow-Labs/DEF-roboticattack)
- **Paper**: arXiv:2411.13587
- **Wave**: 8 (Defense-Only)

## License

Apache 2.0 β€” Robot Flow Labs / AIFLOW LABS LIMITED