Digest | Garden
arrow_back

Articles

Articles 143

Articles

I wrote a thing about this.

Sep 17

Articles

Nearly half a year of silence. We spent it studying one problem: how far RL can scale.

Sep 16

Articles

Just before joining Anthropic, @benkuhn wrote a great post about learning to think independently: ht

Sep 16

Articles

So yes, after a very long time, the article I had been talking about that Im writing since March, is

Sep 15

Articles

A coworker with whom I was discussing inferencing and showed interest that I wanna start learning it

Sep 13

Articles

new post! I argue that current alignment techniques might soon become obsolete (as we scale RL), and

Sep 14

Articles

My 1st simple attempt at an "intuitive RL" blog series. Principles catering to a specific reader (e.

Sep 12

Articles

If you are starting to move past sft into rl-style post-training, these are two really good resource

Aug 26

Articles

A fun small win for automated research:

Aug 27

Articles

Guys and Girl nerds, the tiny 500m-1.2B models trained for a single task are taking off cause they a

Sep 4

Articles

I think this is the craziest thing I've ever read.

Aug 30

Articles

@human_named_gui: had to try this.

Aug 28

Articles

A few days ago, Anthropic shared this brilliant prompt.

Sep 7

Articles

Anyone can simulate the future. But the simulation only matters if it’s trustworthy.

Aug 25

Articles

training and inference communication collectives can be very different (though they use the same col

Aug 23

Articles

Continuous self-improvement needs an ever-expanding supply of training environments (goals).

Aug 20

Articles

I’m mentoring again at MATS this winter. The way I select and supervise MATS fellows is very differe

Aug 21

Articles

almost a year ago, me, @scriptosis and @harish20205 were sitting on a train reading the EAGLE-3 pape

Aug 17

Articles

Weekend hack: Ejipura Flyover tracker 👇

Apr 12

Articles

Experimenting with pi as a minimal app-specific AI interface

Apr 12

Articles

Everyone needs to understand that China’s high-tech push began long before Made in China 2025.

Apr 12

Articles

The clearest article about Anthropic’s work to prevent reward hacking in RL (Chinese): https://t.co/

Apr 12

Articles

very high quality and detailed content if you want to learn how to post train llm!! https://t.co/9mD

Apr 14

Articles

I didn’t used to have friends btw, had to learn how to meet people like this from first principles

Apr 15
143 items
← Prev Next →
auto_stories

Select a bookmark to read

Click any item from the list