Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Future Blog Post
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Blog Post number 4
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 3
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 2
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 1
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
publications
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Published in International Conference on Machine Learning 2024 (Oral), 2024
A benchmark for assessing multimodal LLM-as-a-judge behavior across scoring, pair comparison, and batch ranking tasks.
Recommended citation: Dongping Chen*, Ruoxi Chen*, Shilin Zhang*, Yaochen Wang*, Yinuo Liu*, Huichi Zhou*, Qihui Zhang*, Pan Zhou, Yao Wan, and Lichao Sun. "MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark." International Conference on Machine Learning, 2024. Oral.
Download Paper
What can LLM tell us about cities?
Published in arXiv preprint arXiv:2411.16791, 2024
A study of how large language models encode and expose knowledge about cities and regions worldwide.
Recommended citation: Zhuoheng Li, Yaochen Wang, Zhixue Song, Yuqi Huang, Rui Bao, Guanjie Zheng, and Zhenhui Jessie Li. "What can LLM tell us about cities?" arXiv preprint arXiv:2411.16791, 2024.
Download Paper
Judge Anything: MLLM as a Judge Across Any Modality
Published in KDD 2025 Datasets and Benchmarks Track (Oral), 2025
TaskAnything and JudgeAnything evaluate multimodal understanding, generation, and judging capabilities across any-to-any modality tasks.
Recommended citation: Shu Pu*, Yaochen Wang*, Dongping Chen, Yuhang Chen, Guohao Wang, Qi Qin, Zhongyi Zhang, Zhiyuan Zhang, Zetong Zhou, Shuang Gong, Yi Gui, Yao Wan, and Philip S Yu. "Judge Anything: MLLM as a Judge Across Any Modality." Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2025. Datasets and Benchmarks Track Oral.
Download Paper
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
Published in 39th Annual AAAI Conference on Artificial Intelligence, 2025
Double-Bench provides a multilingual and multimodal benchmark for fine-grained evaluation of document RAG systems.
Recommended citation: Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen, Junjie Yang, Yao Wan, and Weiwei Lin. "Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?" 39th Annual AAAI Conference on Artificial Intelligence, 2025.
Download Paper
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
Published in International Conference on Machine Learning 2026, 2025
A proactive defense framework for LLM agent memory using consensus-based validation and a dual-memory structure.
Recommended citation: Qianshan Wei*, Tengchao Yang*, Yaochen Wang*, Xinfeng Li, Lijun Li, Zhenfei Yin, Yi Zhan, Thorsten Holz, Zhiqiang Lin, and XiaoFeng Wang. "A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory." International Conference on Machine Learning, 2026.
Download Paper
talks
Talk 1 on Relevant Topic in Your Field
Published:
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
Conference Proceeding talk 3 on Relevant Topic in Your Field
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Teaching experience 2
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.