I was watching this , Grant gives a rather interesting example of the way you think about llms. Of how intelligence is the underlying substrate and the auto-regression in a method/interface for interaction with this intelligence.

Imagine an intelligent person locked in box and handed slips of paper each turn to then predict the next work that possibly follows and then your memory is wiped! Now repeat it a bunch of times and at the end to find that someone has made you write an essay through these slips! “You’re a slave to your context(slips).” Imagine you were handed the task of writing the essay directly , you’d have done it/ thought about / written it very differently that just think about one word at a time.

Put in that way(from the video) , it does like a suboptimal interaction mode. “Maybe we have a intelligence locked in the box but auto-regression is a weird way of interacting with it ?”

“Do you get any fruit by questioning the way tokens are generated every now and then?”

I’m not entierly sure but this sounds like “Multi-decoding” or “Speculative Decoding”, have not yet deepdived into these concepts. But from what i think they should be is that, there is someway to train an llm to predict many tokens at ones, making it even think/auto-regress in bigger chunks ( can be larger tokens only , some form of meta network that takes in “logits” of a AR network and then predicts the token-sequence instead of a single token? ) Again going back to man in a box example and drawing parallels with “recursive & hierarchical models” , the locked intelligent man would likely start narrating the essay in larger pieces ones he got a whiff of the task being “essay writing” and would not resort to replying to one slip at a time.


Talking about interfaces and interactions with/of AI. I think there are certainly better UX’s to talk to AI which lets us pick the right interface/mode for a give idea rather that just a plain chat. For example, an exploratory hypothesis is better suited with “why // why-not framing” rather than chat(which can sometimes but not always bring this approach), any method of parallelizing and deliberately isolating contexts and hence outcomes can be interesting. I wonder if there a re other mode’s apart from “y/y-not” that prove to be useful to a “thinker” mode of users.

Another mode can be for writing, llm’s arent that good are producing “great” writing even today, maybe instead of iteratively improving a block of text in context, we wipe the context and seed it everytime with the latest version.

Systematically survey a range of heuristics which could probably provide the right base assumptions required to solve a given problem/task.

Maybe this “range of heuristics” is “diffusion-able in JEPA space” ? Lol, not sure how exactly that could translate but feels interesting.


Maybe every task is a harness away?