LLMs · Reasoning · Agents

Wenxiang Jiao 焦文祥

I study how LLMs and agents learn from data and interaction, reason deeply, and evolve — from capability learning to deep reasoning to self-evolving agents. Recent focus: tool and skill learning, Agentic RL, and Harness Engineering.

Wenxiang Jiao
CurrentXiaohongshu Inc., LLM Algorithm Expert, 2025 - Present
PreviouslyTencent AI Lab, Senior Researcher, 2021 - 2025
EducationPh.D., The Chinese University of Hong Kong, 2021
Service Reviewer for Nature Nature Machine Intelligence NeurIPS ICML ACL

Three papers accepted to NeurIPS 2026, namely MMSkills, Agent2World, and VisFactor. Congratulations!

Three papers accepted to ACL 2026 conference (2 main + 1 findings). Congratulations!

Two papers (REA-RL, DeepCompress) accepted to ICLR 2026. Congratulations!

DeepAgent accepted to WWW 2026. Congratulations!

ArXiv 2026

Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

Trains a 9B harness engineer via RL to patch runtime harnesses from agent failure trajectories, lifting task success from 44.3% to 53.6%.

ArXiv 2026

OmniGAIA: Towards Native Omni-Modal AI Agents

Benchmark and foundation agent for omni-modal AI assistants — multi-hop queries across video, audio, and image, with OmniAtlas trained via hindsight-guided tree exploration and OmniDPO.

NeurIPS 2026

MMSkills: Towards Multimodal Skills for General Visual Agents

Reusable multimodal procedures encoded as state-conditioned skill packages, generated from public trajectories and consulted by a branch-loaded skill agent at runtime.

WWW 2026

DeepAgent: A General Reasoning Agent with Scalable Toolsets

A reasoning agent that tackles general tasks by searching for and using appropriate tools from over 16,000 RapidAPIs in an end-to-end agentic reasoning process.

EMNLP 2024

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

The MAD framework addresses the Degeneration-of-Thought problem in self-reflection and explores divergent chains of thought through structured agent interaction.

ICLR 2024

GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Examines whether safety alignment generalizes to non-natural languages such as ciphers; GPT-4 understands ciphers well enough to produce unsafe outputs.

ICLR 2024 Oral

On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs

Evaluates diverse psychological aspects of LLMs, including personality traits, interpersonal relationships, motivational tests, and emotional abilities.

ArXiv 2023

Is ChatGPT A Good Translator? A Preliminary Study

ChatGPT performs competitively with commercial translation systems on high-resource European languages; GPT-4 further bridges the gap for low-resource and distant languages.

General Agents

Multimodal, tool-using, and collaborative agents for complex tasks.

DeepAgent · OmniGAIA · MAD

LLM Reasoning

Mathematical, reflective, and efficient long-chain reasoning.

DeepCompress · REA-RL

LLM Safety

Risk awareness, jailbreak robustness, multilingual safety, and refusal.

CipherChat · DeRTa

LLM Personality

Emotion, personality, and psychological portrayals in conversational AI.

PsychoBench · EmotionBench · Fints

🌎 Visitor Footprints

... visitors from around the world

Loading visitor map…