Full-manuscript figure quality

SciFigQual-Bench

A benchmark for assessing published CS scientific figures with caption and citing-paragraph context— not isolated crops. Five orthogonal scores on a unified 1–10 scale, plus an auditable staged judge: SFQ-Agent.

ACL · EMNLP · ICML · NeurIPS 2020–2025 Tri-modal: image · caption · context
From isolated-figure evaluation to full-manuscript context binding
5-dim VC · SL · CC · CTX · MR
Protocol 1

Direct

Single VLM pass over figure + caption + citing text.

Protocol 2

Sidecar

Same call, plus OCR/CV side features for denser visual cues.

Protocol 3 · Best

SFQ-Agent

Staged vision / language / fusion + deterministic Runner.

Benchmark at a glance

Scale that matches top-conference figure assessment—with gold labels bound to real manuscript evidence.

0
Curated figures
0
Human gold instances
0
Public test split eval1200
0
Citing-paragraph records

Five-dimensional rubric

Each dimension is scored on [1, 10] with L1 evidence gating: missing caption hides CC; missing citing text hides CTX—never penalize absent metadata as “low quality”.

VC

Visual Clarity

Legibility of text, marks, and encoding at publication scale.

SL

Structure & Layout

Panel organization, reading path, and chartjunk control.

CC

Caption Consistency

Does the caption match visible objects, metrics, and trends?

CTX

Context Consistency

Do citing paragraphs claim only what the figure supports?

MR

Misleading Risk

Higher = lower risk: axes, baselines, unfair comparisons.

From PDFs to gold labels

A deterministic construction funnel: acquire → extract → bind context → annotate → package eval1200.

SciFigQual-Bench construction pipeline

Fig. 2 — Construction pipeline (click to enlarge)

Construction pipeline

62,694 raw PDFs from OpenReview, ACL Anthology, and PMLR are curated into 7,609 figures with index-driven citing paragraphs.

  • Structure-aware extraction (Marker + curation S1–S5)
  • Figure-index context binding (~355k records)
  • Expert five-dimensional annotation + adjudication

Corpus statistics

Venue–topic coverage, domain mix (NLP / ML / CV), temporal span 2020–2025, and per-dimension means. Caption consistency remains the weakest axis—motivating manuscript-grounded CC/CTX evaluation.

  • Mean overall ≈ 8.05 on rated figures
  • Balanced ACL / EMNLP / ICML / NeurIPS splits
  • eval1200: 300 figures × 4 venues
Dataset statistics of SciFigQual-Bench

Fig. 4 — Dataset statistics (click to enlarge)

SFQ-Agent: evidence before scores

Monolithic VLMs conflate perception with text verification. SFQ-Agent separates modalities, then fuses only where needed.

SFQ-Agent scoring pipeline

Fig. 3 — SFQ-Agent staged judging (click to enlarge)

Staged judging

  • Vision locks VC / SL from the figure
  • Language extracts claims from caption & citing text (no pixels)
  • Fusion scores CC / CTX / MR from both evidence reports
  • Runner applies ownership, caps, and deterministic MR mapping

Compared against Direct and Sidecar (+ OCR) on identical inputs.

eval1200 leaderboard highlight

29 protocol–backend configurations. Best overall: SFQ-Agent (F3) with GPT-5.6-Sol— lowest MAE and highest within-±1 agreement versus human gold. Gains come from protocol–evidence alignment, not model scale alone.

0
Overall MAE ↓ (F3)
0%
Within-1 agreement ↑
0
Spearman SRCC ↑

Citation

If you use SciFigQual-Bench or SFQ-Agent, please cite:

@article{scifigqual2026,
  title={SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context},
  author={Deng, Zihan and Xu, Chuanzhi and Liang, Huiqi and Li, Haoyang and Zhong, Xiaozhen and Yu, Lequan},
  year={2026}
}