Today is my first day at Meta Superintelligence Labs. I’ll be focusing on alignment and safety, building on my time at Scale Research and SEAL. Grateful to keep working with —no one more committed, clear-eyed, or mission-driven. Excited for what’s ahead 
Summer Yue
Summer Yue
117 posts
Summer Yue
@summeryue0
Safety and alignment at Meta Superintelligence. Prev: VP of Research at Scale AI, research at Google DeepMind / Brain (Gemini, LaMDA, RL / TFAgents, AlphaChip).
San Francisco, CA
Summer Yue’s posts
I’m joining Scale and we are starting a new safety lab! Hiring researchers interested in trustworthy evaluations, red teaming and scalable oversight. These areas require hands-on interaction with human data, and Scale is an unparalleled place to do it.
1. Claude 3.5 Sonnet is now #1 in Instruction Following on the SEAL leaderboards (scale.com/leaderboard) 
Excited to share our latest research on red teaming and agent safety from SEAL team at .
This work highlights a critical gap: safety mechanisms in advanced LLMs do not generalize well to downstream browser agents. We also found that LLM attacks transfer with high
Quote
Zifan (Sail) Wang
@_zifan_wang
(1/7) Excited to share our new red teaming work at Scale, Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents. We find jailbreaking LLM agents that use browsers is surprisingly easy. In many cases, you can just direct ask!
Paper & Project page: scale.com/research/brows
How much do LLMs overfit public benchmarks? Our team at SEAL lab studied this by creating a GSM8k-equivalent eval from scratch. The resulting performance gap reveals data contamination in some model families, while GPT, Claude, and Gemini show no signs of overfitting.
Quote
Announcing our latest SEAL Leaderboard on Adversarial Robustness!
Red team-generated prompts
Focused on universal harm scenarios
Transparent evaluation methods
SEAL evals are private, expert evals that refresh periodically: scale.com/leaderboard
LLMs are often evaluated against single-turn automated attacks. This is an insufficient threat model for real-world malicious use, where malicious humans chat with LLMs over multiple turns.
We show that LLM defenses are much less robust than the reported numbers suggest.
SEAL Visual-Understanding Leaderboard Launch
Today, we’re introducing VISTA—a new rubric-based visual task assessment benchmark that pushes beyond simple Q&A.
The leading models achieve under 40% on this eval, compared to a human baseline of ~55.4%. This highlights that
1.
Exciting update: Claude 3.5 Sonnet is now #1 in Coding on the SEAL leaderboard (scale.com/leaderboard)! 
Quote
Alexandr Wang
@alexandr_wang
As LLMs get smarter, evals need to get harder.
OpenAI’s o1 has already maxed out most major benchmarks.
Scale is partnering with CAIS to launch Humanity’s Last Exam: the toughest open-source benchmark for LLMs.
We're putting up $500K in prizes for the best questions.
(read on)
Quote
Sundar Pichai
@sundarpichai
We're expanding access to Bard in US + UK with more countries ahead, it's an early experiment that lets you collaborate with generative AI. Hope Bard sparks more creativity and curiosity, and will get better with feedback. Sign up: bard.google.com
blog.google/technology/ai/
If a model lies when pressured—it’s not ready for AGI.
The new MASK leaderboard is live.
Built on the private split of our open-source honesty benchmark (w/ ), it tests whether models lie under pressure—even when they know better.
Leaderboard:
Do LLMs hold knowledge that might be dangerous in the hands of a malicious user? Can hazardous knowledge be unlearned?
Introducing WMDP: an open-source eval benchmark of 4,157 multiple-choice questions that serve as a proxy measurement of LLM’s risky knowledge in biosecurity,
Excited to share our latest work “Jailbreaking to Jailbreak (J2)”, from the SEAL team and 's Red Team! As frontier models become more creative and capable of reasoning, they can now not only assist human red teamers but also autonomously drive red teaming efforts.
Will be at #NeurIPS2023 next week! If you’re an LLM researcher / research engineer interested in robust evaluations, safety, red teaming or scalable oversight, let’s chat! Mainly hiring for SEAL but also happy to chat about collaboration opportunities.
Can robust LLM defenses be jailbroken by humans?
We show that Scale Red teamers successfully break defenses on 70+% of harmful behaviors, while most automated adversarial attacks yield single-digit success rates. 
Gemini is out and 90%+ MMLU! Huge congrats to my friends and former colleagues and everyone who were part of this achievement. Truly fantastic team work!
Quote