#Benchmarking

26 статей
arXiv Computer Vision
arXiv Computer Vision

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

arXiv:2608.29607v1 Announce Type: new Abstract: Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the mos...

arXiv Computer Vision
arXiv NLP
arXiv NLP

PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating Text-to-Text Privatization

arXiv:2608.29624v1 Announce Type: new Abstract: Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text pr...

arXiv NLP
arXiv NLP
arXiv NLP

SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training

arXiv:2608.29481v1 Announce Type: new Abstract: Pediatric serious illness communication (SIC) is critically important, yet scalable communication training for clinicians remains limited. Compared with...

arXiv NLP
arXiv Computer Vision
arXiv Computer Vision

SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models

arXiv:2605.31597v4 Аннотация Тип: замена Абстракции: Измерение понимания структурированных объектов в моделях построения видения остается сложной задачей из-за непоследовательных протоколов оценки и о...

arXiv Computer Vision
arXiv Computer Vision
arXiv Computer Vision

ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification

arXiv:2608.27150v1 Анонс Тип: кросс-абстракт: Классификация объектов в случае компьютерного зрения на основе событий является задачей, которая привлекает значительное внимание исследователей. Классифи...

arXiv Computer Vision
arXiv Machine Learning
arXiv Machine Learning

Benchmarking Fast Domain Adaptation for Unsupervised Speech Units

arXiv:2608.26992v1 Анонс Тип: новый Абстракт: Презентативное обучение привлекло большое внимание и сумело достичь хороших результатов в качестве метода предварительной подготовки для задач ниже по теч...

arXiv Machine Learning
1
Читать
arXiv Computer Vision
arXiv Computer Vision

CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

arXiv:2608.15404v2 Анонс Тип: замена Абстракт: Концептуальные модели Боттлнека (CBM) предназначены для визуальной интерпретации классификации путем выражения прогнозов с помощью понятных человеку конц...

arXiv Computer Vision
arXiv AI
arXiv AI

Оригинальное название: The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

arXiv:2608.22301v1 Человечество имитирует на уровне намерений: давая демонстрацию, мы выводим ее цель и выполняем ее с помощью любых инструментов, объектов и макетов. Современные политики роботов вмес...

arXiv AI
1
Читать
arXiv NLP
arXiv NLP

Оригинальное название: Hear2Act: Benchmarking When Prosody Should Change Что делает помощник

arXiv:2608.19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves re...

arXiv NLP
1
Читать
arXiv AI
arXiv AI

FinRCA-Bench: поиск доказательств и обоснование финансовых систем ИИ

arXiv:2608.18534v1 Большие языковые модели все чаще используются для поддержки финансовых операций, но их кажущаяся производительность рассуждений может зависеть от того, получают ли они правильные до...

arXiv AI
1
Читать
arXiv Machine Learning
arXiv Machine Learning

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

arXiv:2608.17293v1 Announce Type: new Abstract: Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Ex...

arXiv Machine Learning
arXiv Machine Learning
arXiv Machine Learning

Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

arXiv:2608.16928v1 Announce Type: new Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification rang...

arXiv Machine Learning
arXiv NLP
arXiv NLP

When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

arXiv:2608.17979v1 Announce Type: new Abstract: Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practic...

arXiv NLP
arXiv Computer Vision
arXiv Computer Vision

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

arXiv:2608.17426v1 Announce Type: new Abstract: We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achieve...

arXiv Computer Vision
arXiv Computer Vision
arXiv Computer Vision

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

arXiv:2608.14976v1 Announce Type: new Abstract: Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests inv...

arXiv Computer Vision