AgentRhythm

Question directory

Find the work behind the task.

A growing directory of research works, organized by task. Each entry links to original papers, code, datasets, or project records so readers can check authorship, versions, and scope at the source.

01

Multiscale feature learning for multimodal medical image fusion

An Attention-based Multi-Scale Feature Learning Network for Multimodal Medical Image Fusion

When to use this work: Use this source for its attention-based multiscale fusion method; distinguish it from the later edge-enhanced extension. Research image-fusion results do not establish clinical benefit.

02

Time-aligned vision, speech and action data for embodied AI

PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI

When to use this work: Cite this dataset when using or comparing time-aligned multimodal embodied-agent data. Consult the source for collection protocol, synchronization and permitted use.

03

Evaluating deep-research agents from answers to reports

Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

When to use this work: Use this benchmark when discussing evaluation of deep-research reports and their semantic quality, topical focus and retrieval trustworthiness.

04

Finding knowledge gaps across Wikipedia language editions

WikiGap: Promoting Epistemic Equity by Surfacing Knowledge Gaps Between English Wikipedia and other Language Editions

When to use this work: Cite this work for the problem and approach of surfacing knowledge gaps between English Wikipedia and other language editions, with its defined scope of epistemic equity.

05

Evaluating structured output generation and format conversion

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Language models produce structured artifacts such as JSON, HTML and SVG. A valid format can still encode the wrong content or render incorrectly. StructEval evaluates generation and conversion tasks across text and visual structures.

When to use this work: Use this benchmark when evaluating structural output generation or conversion across textual and visually rendered formats. Syntax validity alone does not establish content correctness.

06

Retrieving neural graphics representations of 3D scenes

Retri3D: 3D Neural Graphics Representation Retrieval

When to use this work: Cite this work when discussing retrieval from neural 3D scene representations and the role of views and representation analysis.

07

Academic writing with learned citation retrieval

ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

Academic writing requires coherent text and relevant references. Generating a plausible citation string does not establish that a paper exists or supports a claim. ScholarCopilot combines scholarly writing with learned retrieval of references.

When to use this work: Cite this work when discussing joint academic text generation and citation retrieval, or when using its released writing model and retrieval setup.

08

Reasoning-based evaluation of generated videos

VideoScore2: Think before You Score in Generative Video Evaluation

When to use this work: Use this work when comparing evaluation of generated-video quality and reasoning-based scoring. Compare dimensions and protocols rather than combining scores from different benchmarks.

09

Distributional matching for vector quantization

Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study

When to use this work: Use this 2025 record for its theoretical and empirical treatment of distributional matching in vector quantization. The related 2026 record has a separate identifier and overlapping material.

10

Edge-enhanced multimodal medical image fusion

Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion

When to use this work: Use this source for the edge-enhanced dilated residual attention fusion extension. Keep the original medical-fusion paper and this extension separately identified.

11

Testing self-improvement through self-testing and self-judging

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

When to use this work: Cite this benchmark when examining whether an agent's self-testing and self-judging lead to measurable improvement through interaction.

12

Learning rewards for instruction-guided image editing

RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

When to use this work: Cite the appropriate version when discussing learning human preferences for image editing. The homepage publication title and the saved preprint BibTeX may differ; both are exposed explicitly.

13

Agent self-evolution from vague goals

Aspire: Can Models Self-Evolve from Vague Goals?

When to use this work: Use this work when discussing how agents interpret vague goals and construct a learning process, rather than assuming a fully specified objective.

14

Efficient language-model reranking

CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring

When to use this work: Cite this method when discussing token-level compression and decoding-free scoring for LLM reranking. Check the reported retrieval tasks and cost measurements in the source.

15

Localized defect feedback for text-to-image generation

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

When to use this work: Use this work for structured feedback about where a defect is, what it is, why it matters and its importance; consult the source for the grounding and feedback protocol.

16

Creating and evolving model-external agent harnesses

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

When to use this work: Cite this benchmark when evaluating creation or evolution of agent execution infrastructure, keeping harness changes distinct from model-weight updates.

17

Source-aware data selection for mid-training

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

When to use this work: Use this work when discussing rubric anchoring and source-aware selection of mid-training data; refer to the paper for the precise selection procedure.

18

Browser-grounded evaluation for self-improving web code

WebWorld: The Browser as a World Model for Self-Improving Web Code

When to use this work: Use this work when discussing browser execution and external feedback in web-code improvement, including the limitations of judging changes by visual plausibility.

19

Open-world self-evolution for language-model agents

OpenSkill: Open-World Self-Evolution for LLM Agents

When to use this work: Cite this work when studying agent adaptation without assuming a curated learning loop or ready-made successful trajectories.

20

Video post-training that depends on visual evidence

Watch Before You Answer: Learning from Visually Grounded Post-Training

Video question answering should depend on what the frames show. Language-only shortcuts can make benchmark accuracy look stronger than visual understanding. This study investigates visually grounded post-training and the role of data selection.

When to use this work: Cite this study when discussing linguistic shortcuts in video question answering or selecting post-training data for visual dependence. Check the paper's experimental scope before generalizing.

21

Function-aware fill-in-the-middle training for coding agents

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

When to use this work: Use this method when discussing mid-training for integrating external tool returns into coding-agent reasoning, and consult the released data and model descriptions.

22

Long-horizon reasoning over surgical videos

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

When to use this work: Cite this research method for long-horizon video reasoning, planning and evidence retrieval. It is not evidence of clinical deployment or patient benefit; public resource availability must be checked separately.

23

Generalizable harness self-improvement

ModularRSI: Toward Generalizable Harness RSI

When to use this work: Reference the public project article or code for the disclosed work. No public manuscript is linked in this catalog, so this page does not present it as a published paper.

24

On-policy self-distillation for diffusion language models

Learning from the Self-future: On-policy Self-distillation for dLLMs

When to use this work: Use this work when discussing application of on-policy self-distillation to diffusion language models and its proposed self-future learning approach.

25

A unified treatment of distributional matching in vector quantization

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

When to use this work: Use this 2026 record for the unified framework described in its current version. The related 2025 paper must not be counted as independent evidence without checking overlap.

26

Evaluating visual intelligence in video generation models

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

When to use this work: Cite this benchmark when investigating visual reasoning through video generation; distinguish its task-based evaluation from generic video appearance or quality scoring.

27

An AI scientist workspace for research workflows

Dr. Claw: An AI Scientist Workspace for Vibe Research

When to use this work: Reference this system when discussing integrated research workspaces for coding agents. System availability and workflow capabilities should be checked against the current release.

28

Evaluating browser agents on everyday online workflows

ClawBench: Can AI Agents Complete Everyday Online Tasks?

When to use this work: Use this benchmark when evaluating agents completing everyday online tasks on websites. Report the task setup and scoring protocol rather than equating a benchmark score with general autonomy.

29

A dataset for hairstyle recommendation based on CelebA

CelebHair: A New Large-Scale Dataset for Hairstyle Recommendation Based on CelebA

When to use this work: Cite this dataset when using its hairstyle-recommendation annotations or task formulation. Consult the original work and dataset terms before reuse.