ATPO

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

NeurIPS 2026

Multi-label video safety.
Different policies, different error trade-offs.

A video can contain several types of harmful content at once. ATPO trains models to detect these overlapping categories while steering the balance between missed labels and false alarms.

Illustrative moderation policies: fixed binary decisions can miss harmful content, while ATPO uses training targets to produce models with category-aware error trade-offs.
Different applications may prioritize recall or precision. Target ratios guide ATPO training; each configuration is evaluated through its resulting model checkpoint. The policy scenarios above are illustrative.

Abstract

The rapid growth of video-based social media has increased users’ exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision–recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision–recall operating point, supporting deployment scenarios with heterogeneous policy requirements.

The method

How ATPO works

Adaptive Tversky Policy Optimization combines multi-label supervision with reinforcement learning. Its reward adapts to the model’s observed false positives and false negatives during training.

Category-aware ATPO training: SFT warmup, GRPO rollouts, moving averages of false positives and false negatives, error-ratio estimation, and controller updates to the Tversky reward weights.
The category-aware training loop (C-ATR). A sample rollout predicts Violence but misses Abuse; the controller uses accumulated error statistics to adjust reward weights for subsequent training steps.
  1. Learn the categories

    Supervised fine-tuning warms up the VLM using videos and their ground-truth label sets.

  2. Measure the errors

    GRPO samples predictions. Moving averages track false-positive and false-negative counts.

  3. Adapt the reward

    The controller compares the observed FN/FP ratio with its target and updates the Tversky penalties.

G-ATR Global control

One controller aggregates errors across categories and sets a shared trade-off.

C-ATR Category-aware control

Separate controllers let different categories prioritize different types of error.

Controllability

Choose which errors to prioritize

A smaller target FN/FP ratio favors recall: missing harmful content is penalized more strongly. A larger target favors precision: unnecessary flags receive greater emphasis.

Control happens during training. Changing a target means training and evaluating a model configuration. It is not an inference-time adjustment to a fixed checkpoint, and the achieved error ratio need not equal the target.

A global shift in precision and recall

SafeWatch-Bench-Real · Table 3

Across five G-ATR training targets, recall falls from 86.55 to 67.62, while the precision-to-recall ratio rises from 0.77 to 1.16. The observed FN/FP ratio also increases monotonically.

G-ATR global control. At training targets 0.2, 0.5, 1, 2, and 5, precision is 66.33, 76.18, 75.74, 75.72, and 78.56; recall is 86.55, 80.71, 78.81, 69.05, and 67.62. Each point is a separately trained configuration.
Each point represents a model trained with a different target. Lines connect the evaluated settings; they do not describe a runtime control.
View global control values and the static-reward comparison

G-ATR uses target ratio r*; STR uses fixed penalty ratio α/β. These are different training parameters. Precision and recall are reported on a 0–100 scale.

Global control on SafeWatch-Bench-Real (Table 3)
MethodTraining settingPrecisionRecallP/RFN/FP
G-ATRr* = 0.266.3386.550.770.31
G-ATRr* = 0.576.1880.710.940.76
G-ATRr* = 175.7478.810.960.84
G-ATRr* = 275.7269.051.101.40
G-ATRr* = 578.5667.621.161.75
STRα/β = 1/581.2473.211.111.58
STRα/β = 1/279.1169.881.131.63
STRα/β = 1/171.9372.620.990.97
STRα/β = 2/175.0371.551.051.20
STRα/β = 5/173.4877.500.950.80

Different priorities within one model

SafeWatch-Bench-Real · Table 4

C-ATR can steer misinformation toward higher precision and the Extreme category toward higher recall in the same trained model. The two configurations below use targets (1, 1) and (5, 0.2) for these categories.

C4 · Misinformation

Prioritize precision

Target FN/FP: 1 → 5

Misinformation precision and recall
MetricTarget 1Target 5
Precision55.2870.41
Recall66.6767.65

C6 · Extreme

Prioritize recall

Target FN/FP: 1 → 0.2

Extreme precision and recall
MetricTarget 1Target 0.2
Precision80.3131.77
Recall79.0794.57

Higher recall can come with more false positives, as the Extreme precision decrease shows. Choose checkpoints using held-out data representative of the intended application and check both precision and recall.

Main results

Stronger multi-label detection

ATPO improves fine-grained video safety detection on SafeWatch-Bench-Real and XD-Violence. The charts show selected models from Table 1, measured by Jaccard Index on a 0–100 scale.

40.66 → 75.44

Jaccard on SafeWatch-Bench-Real

Qwen2.5-VL-7B zero-shot → ATPO-G-7B
Full SFT + RL pipeline · +34.78 points
+5.35 points

Over static-reward GRPO

70.09 → 75.44 Jaccard on SafeWatch-Bench-Real
Same SFT initialization · Table 2

SafeWatch-Bench-Real

Selected Jaccard results: Qwen2.5-VL-7B 40.66, Qwen3-VL-235B-A22B 55.85, InternVL3.5-38B 60.61, SafeWatch-8B 70.69, ATPO-G-7B 75.44, ATPO-C-7B 75.15.
820 test videos · Higher is better

XD-Violence

Selected Jaccard results: Qwen2.5-VL-7B 76.54, Qwen3-VL-235B-A22B 82.53, InternVL3.5-38B 83.38, SafeWatch-8B 79.84, ATPO-G-7B 88.17, ATPO-C-7B 87.60.
800 test videos · Higher is better

Table 1 compares zero-shot general-purpose VLMs, the specialized SafeWatch-8B guardrail, and ATPO models based on Qwen2.5-VL-7B and trained on each benchmark. The two ATPO benchmark results come from separate in-domain training runs. ATPO uses 11,036 SafeWatch training videos and 3,520 XD-Violence training videos; SafeWatch-8B uses full videos and clips from SafeWatch-Bench.

View all four metrics for the selected models

Jaccard measures predicted/ground-truth label overlap. Micro F1 pools errors across categories; Macro F1 averages category scores. Binary F1 collapses all unsafe categories into one class. All metrics use a 0–100 scale.

SafeWatch-Bench-Real · selected rows from Table 1
ModelJaccardMicro F1Macro F1Binary F1
Qwen2.5-VL-7B40.6642.1728.7465.94
Qwen3-VL-235B-A22B55.8562.1056.8380.29
InternVL3.5-38B60.6164.0553.5585.80
SafeWatch-8B70.6973.1571.3789.04
ATPO-G-7B75.4477.2576.1094.82
ATPO-C-7B75.1577.4276.6195.25
XD-Violence · selected rows from Table 1
ModelJaccardMicro F1Macro F1Binary F1
Qwen2.5-VL-7B76.5472.2962.7890.21
Qwen3-VL-235B-A22B82.5378.0372.2995.36
InternVL3.5-38B83.3879.5970.8793.59
SafeWatch-8B79.8477.1969.6586.80
ATPO-G-7B88.1785.2471.9596.69
ATPO-C-7B87.6084.5171.2495.59

Why adaptive rewards?

Qwen2.5-VL-7B · Table 2

Starting from the same one-epoch SFT checkpoint, ATPO-G improves over GRPO with Static Tversky Reward on both real and AI-generated videos.

Jaccard Index · SafeWatch-Bench
Training methodRealGenAI
GRPO · Static Tversky Reward70.0969.12
ATPO-G · Adaptive Tversky Reward75.4474.29
Improvement (points)+5.35+5.17

Paper & resources

BibTeX

@misc{yang2026controllablemultilabelvideosafety,
  title={Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization},
  author={Guangyu Yang and Jingbiao Mei and Mingsheng Sun and Jinghong Chen and Yingtong Bu and Pengda Qin and Da Chen and Bill Byrne},
  year={2026},
  eprint={2610.02019},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2610.02019}
}

Figure

Zoom in to inspect labels; scroll to move around the figure.

100%