Phones don’t have lights

Mark Zuckerberg has a new defense of the Ray-Ban Meta glasses: they’re actually doing more to signal they’re taking a photo than phones do. He’s brought this up in at least two recent interviews, noting that the glasses have a light that comes on to signal when a photo is being taken, while phones do … Read more

Tesla’s Optimus robot is going through growing pains

Hitting its goal of making 20,000 Optimus robots per week is reportedly proving tricky for Tesla. The Information reports that Tesla produced “several hundred robots a week” last month, after it repurposed its Model S and Model X production lines for Optimus earlier this year. However, this strategy is reportedly creating manufacturing snags, like issues … Read more

Meta makes the Muse filesystem even more accessible

Yesterday, with a little prodding, it was discovered that Meta’s Muse would expose its filesystem to curious users. The files offered a fascinating peek under the hood of an AI chatbot, and appeared to expose details we weren’t meant to see, not least because Muse itself told people, including us, it wasn’t supposed to reveal … Read more

Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

When you post-train a Mixture-of-Experts (MoE) model with Reinforcement Learning from Human Feedback (RLHF) or Group Relative Policy Optimization (GRPO) at scale, three simultaneous challenges emerge. The first requires coordinating heterogeneous compute for rollout generation and policy training. Second, sustaining high-throughput communication across hundreds of accelerators. And third, dynamically orchestrating every subsystem to keep them … Read more

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Reinforcement learning (RL) post-training is becoming a standard step in building capable language model agents. Models learn to reason and act across sequences of steps by generating trajectories, receiving rewards, and updating their policy based on outcomes. Running this at scale, across multiple nodes with hundreds of GPU-hours of rollouts per training run, requires persistent … Read more