> Applied statistics is a
far more precise descriptor, “but no one wants to use that term, because
it’s not as sexy.”
This really hit me some time back when I was explaining AI to a friend. After about 10 mins of rambling about LLMs and mentioning the attention paper like I knew what I was talking about, it ended with “oh so it’s just a really advanced auto correct”
As an MLE I feel these takes are too reductionist.
You could say the (nearly) same thing about search. And content recommendation. And clustering. And topic modeling. And outlier detection. And spam filtering. And image diffusion. And dimension reduction. And...
There's a lot in common between these things, but there's also a lot cool and different!
For transformers in particular, it's pretty cool that you get some WILD emergent properties simply from scaling up.
So yes, it's just a next token predictor, but I'm just a bundle of nerves and meat. I don't get a lot out of those descriptions.
It wouldn't require perfect accuracy, just rough accuracy and a human to confirm, and it's already more than good enough for that. I do not understand this confusion surrounding modern math.
- Many more mediocre papers written (mediocre ideas, implementation, claude-isms everywhere)
- Much easier to try every possible combination of a regression in order to show the result you want (same for theorists).
The one thing I'm happy about is it's now much easier to extract historical data from old documents from Google Books. Still not perfect, but takes you 95% there. And creating plots and datavis just for quick exploration is super fast.
Here's the thing about economists... The loudest ones don't want to be correct, they want to be influential. The ones who can actually make good predictions work for banks and hedge funds lol.
This really hit me some time back when I was explaining AI to a friend. After about 10 mins of rambling about LLMs and mentioning the attention paper like I knew what I was talking about, it ended with “oh so it’s just a really advanced auto correct”
You could say the (nearly) same thing about search. And content recommendation. And clustering. And topic modeling. And outlier detection. And spam filtering. And image diffusion. And dimension reduction. And...
There's a lot in common between these things, but there's also a lot cool and different!
For transformers in particular, it's pretty cool that you get some WILD emergent properties simply from scaling up.
So yes, it's just a next token predictor, but I'm just a bundle of nerves and meat. I don't get a lot out of those descriptions.
- Many more mediocre papers written (mediocre ideas, implementation, claude-isms everywhere)
- Much easier to try every possible combination of a regression in order to show the result you want (same for theorists).
The one thing I'm happy about is it's now much easier to extract historical data from old documents from Google Books. Still not perfect, but takes you 95% there. And creating plots and datavis just for quick exploration is super fast.