Menu

Research

SoVTR: Set-of-Vision-Text Rewards for Attribute-Grounded Emotion Recognition in VLLMs

Supervised fine-tuning penalizes every reasoning token uniformly, which suppresses valid alternatives and erodes reasoning diversity. SoVTR is a reward-based framework that decomposes supervision into a Face Reward (action units, gaze, micro-expressions), a Body Reward (pose and gesture), and a Context Reward (person-object relations and scene activity), combining BERTScore attribute matching with GRPO policy optimization. It improves both accuracy and grounding quality on DFEW, FERV39k, and ExpW.

Read More

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

A story agent, an image agent, and a critic agent collaborate to fabricate social media posts that plausibly counter true news. Applying the framework yields over 9,000 paired multimodal news posts across science, health, and entertainment, against which 16 open- and closed-source MLLMs are benchmarked for detection. Most fall well short of human accuracy and fail critically at judging image authenticity.

Read More

Toward Federated Large Language Models in Medicine: A Parameter-Efficient Framework for Privacy-Preserving, Multi-Institutional Adaptation

Clinical text is siloed inside individual institutions, which blocks the centralized fine-tuning that medical language models normally depend on. This work pairs federated learning with parameter-efficient adaptation so each site trains locally and shares only a small set of updated parameters, never raw patient records.

Read More
Hidden Visitor Tracker Hidden Visitor Counter