Digest | Garden
All Bookmarks 3319

Threads

A reward can tell an RL/online learner that something worked without telling it which combination of

Aug 20

Videos

Want to watch a 535B parameter (23B active) LLM get trained live? Follow along here https://t.co/ELz

Aug 20

Articles

Continuous self-improvement needs an ever-expanding supply of training environments (goals).

Aug 20

Threads

@askalphaxiv: "Recirculation"

Aug 19

Threads

absolute GPT-3 moment for robotics

Aug 19

Papers

Very interesting new work from Microsoft.

Aug 19

Code

Rocket One has introduced Swarm Stage AI, a platform designed to simulate coordinated drone-swarm th

Aug 19

Papers

What seminal papers should every grad student in deep learning know?

Aug 19

Books

Great list of NLP papers/code to read by @Harvard https://t.co/FYIY4iBgtp

Aug 18

Papers

introducing the @tastelabs research fellowship!

Aug 18

Threads

Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models.

Aug 18

Threads

Caddy's your new personal trainer

Aug 18

Threads

On KDA prefill (Kimi K3's attention), llms independently writes kernel that gets a 2.05x speedup ove

Aug 18

Papers

This Google DeepMind paper is f*cking brilliant

Aug 18

Products

Most Transformer interview questions sound easy until someone asks you to explain what’s actually ha

Aug 17

Books

RL for LLMs famously provides just 1 bit per rollout as learning signal, much less than SFT's dense

Aug 17

Articles

almost a year ago, me, @scriptosis and @harish20205 were sitting on a train reading the EAGLE-3 pape

Aug 17

Code

someone built a repo "powered by netherite" that uses VLMs with super fast distributed RL to teach m

Aug 17

Papers

SimpleOPD transfers a long-context teacher's proof behavior across tokenizer boundaries by aligning

Aug 17

Papers

Reinforcement learning with verifiable reward (RLVR) is the technique behind the recent incredible b

Aug 17

Papers

New Meta paper shows, small models may not be bad predictors of scale; they may just be getting unde

Aug 16

Articles

NVIDIA has lost it..

Aug 16

Threads

over the past couple weeks, i've gotten a primitive form of continual in-context learning to work de

Aug 16

Threads

Latest Deep RL class lectures are now online!

Aug 15
auto_stories

Select a bookmark to read

Click any item from the list