Pengfei Liu

500 posts
Opens profile photo
Pengfei Liu
@stefan_fee
Associate Prof. at SJTU, leading GAIR Lab (plms.ai) Co-founder of Inspired Cognition, Postdoc at , Previously FNLP, ,

Pengfei Liu’s posts

Pinned
Seedance 2.0 is impressive. But it's closed-source! Introducing our daVinci-MagiHuman — a single-stream 15B Transformer trained from scratch that jointly generates video + audio. No cross-attention. No multi-stream branches. Just self-attention. ⚡ 5s 1080p video in 38s on a
2:00
What is prompt-based learning, and what challenges are there? Will it be a new paradigm or a way for human-PLMs communication? How does it connect with other research and how to position it in the evolution of the NLP research paradigm? We released a systematic survey and beyond
Image
Image
Image
What's your system good/bad at? Where can your model outperform others? What are the mistakes that the top-10 systems make? We are always struggling with these questions. A new academic tool can help us answer them in a one-click fashion and many more:explainaboard.nlpedia.ai
GIF
Crazy finding!!!!! -> ” Without introducing any additional data or advanced training techniques, and merely by reformatting the response, LLaMA-2-13B’s mathematical reasoning ability on GSM8K can be improved from 46.77% to 56.63% in accuracy"
Image
Quote
Run-Ze Fan
@Vfrz525_
Been diving into some papers on data synthesis lately, especially those about enhancing math reasoning. Most of them seem to miss our work on 'Reformatted Alignment' (arxiv.org/abs/2402.12219)—another approach to boosting data for math reasoning.
We present a new paradigm for evaluating generated text, BARTScore: conceptualize evaluation as a text generation problem. We empirically demonstrated benefits using such framework: multi-perspective evaluation, prompt, and winning in 16 of 22 settings. Demo/Code/Paper (1/n)
Image
We have one paper. ACL20 reviewer: this should be a long paper, reject. After extending it to a long version, EMNLP20 reviewer: this should be a short paper, reject... 😉
Quote
Yonatan Belinkov
@boknilev
Short paper acceptance rate is very low, which seems to continue a trend in recent #NLProc conferences. Are reviewers expecting too much from short papers? x.com/emnlp2020/stat…
How much data *leakage* do popular LLMs have on public benchmarks? Our recent work will tell you: Benchmarking Benchmark Leakage in Large Language Models arxiv.org/pdf/2404.18824
Image
Quote
Alexandr Wang
@alexandr_wang
Image
How overfit are popular LLMs on public benchmarks? New research out of @scale_ai SEAL to answer this: - produced a new eval GSM1k - evaluated public LLMs for overfitting on GSM8k VERDICT: Mistral & Phi are overfitting benchmarks, while GPT, Claude, Gemini, and Llama are not.
The last step of the #ACL2021 submission: manual text compression! We should keep our data before and after compression so that we can automate this process next year😂
Image
OpenAI o1-preview achieved an average improvement of 20+ points on OlympicArena (e.g., math, physics, chemistry, bio, geo, astro) !!! However, there's still significant room for improvement (average score below 60 points)
Quote
Zhen Huang
@Z_Huang_02
🔥o1-preview has shown incredible improvements in reasoning ability across complex disciplines on our OlympicArena (val + text-only) subset! We’re also eagerly looking forward to the performance of the multimodal version of o1 in the future! x.com/Z_Huang_02/sta…
Image
Image
I'm super honored that our paper was awarded as "Best Demo Paper" at ACL 2021!! Special thanks to all awesome collaborators, reviewers' insightful suggestions, and the committee's recognition. We will continue optimizing ExplainaBoard to make it a truly useful tool!
Image
Quote
Pengfei Liu
@stefan_fee
Embedded video
GIF
What's your system good/bad at? Where can your model outperform others? What are the mistakes that the top-10 systems make? We are always struggling with these questions. A new academic tool can help us answer them in a one-click fashion and many more:explainaboard.nlpedia.ai
Three months on, let's look back and see where the field of prompt-based learning is at? 74 -> 120+ papers in the past two months
Image
Quote
Pengfei Liu
@stefan_fee
Image
Image
What is prompt-based learning, and what challenges are there? Will it be a new paradigm or a way for human-PLMs communication? How does it connect with other research and how to position it in the evolution of the NLP research paradigm? We released a systematic survey and beyond
Just as prompts can connect human and large language models, Inspired Cognition aims to close human with AI by standardizing chaos, integrating fragmentation, and democratizing domain expertise:)
Quote
Graham Neubig
@gneubig
Happy to announce that I've formed a company, Inspired Cognition (inspiredco.ai) together with @stefan_fee and @odashi_en! Our goal is to make it easier and more efficient to build AI systems (particularly NLP) through our tools and expertise. 1/2
Image