#Benchmarking

26 статей
arXiv Computer Vision
arXiv Computer Vision

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

arXiv:2608.14778v1 Announce Type: new Abstract: Hepatocellular carcinoma (HCC) is the third leading cause of cancer-related mortality worldwide, with early detection improving survival from 70\%. The ...

arXiv Computer Vision
arXiv Quantitative Biology
arXiv Quantitative Biology

Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding

arXiv:2605.00865v3 Announce Type: replace-cross Abstract: We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, predict...

arXiv Quantitative Biology
arXiv AI
arXiv AI

Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding

arXiv:2608.14179v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for percep...

arXiv AI