The more and more LLMs are getting better, the more parallels I’m start to see the parallels between “predict next token” and how I myself think. It’s not intelligent, there is no doubt about it, but are “we”? Am I “intelligent”? I mean sure I would believe I have a sense of judgement that it does not have, so there has to be someone guiding it; but the same can be said of working with other people. The “most likely” next token reward hack of course will not end up making something very creative, after all my taste is the sum total of all the paths that I had to take in life, there was never a million GB dataset and billions of dollars spent on me to make my paths as globally optimal as possible; an optimality that sits at the “mean”.

Think about how we do reasoning. I remember when I was younger, I spent a some time trying to be faster at doing mental math — the usual arithemetic operations. Everyone has at some point memorised multiplication tables, as part of a long multiplication, to recall what 8x7 is I in my head did verbalised it, and that was one of the steps that I tried to speed up on. First optimisation was just verbalising digits , so it was 8 7 56, next was skipping the verbalisation altother and just verbalising the finished results. So something like

23 x 45 became 15(5x3) 5(digit) 12(3x4) 10(5x2) 22(add) 23(+1 carry) 8(2x4) 10(+2 carry) => 10 3 5
23
45

And now I see the same behaviors in an LLm, emerging organically. The way CoT works is entirely this reasoning that I too needed to do ( for there simply is no other basis of directly getting to a result with reasoning like a language ), second one being skipping intermediate tokens or using a token dense language like Chinese for thinking.

Another is context, that is where we without a doubt ( for now atleast ), but we still lose in terms of recall so I guess it’s better to call it a tradeoff. There is an almost perfect recall for anything that has fit within an LLM’s context window, obviously as it has the complete representation for it. For me, well by all means I cannot recite from memory the entire code file of a 1000 lines, that is just not possible but I have do have an incredibly long range sparse context, like I can immediately and without any concious effort ( apart from reading/glacing at the code ) extract the high level structure of it, like what is where, and then work with that knowledge; akin to some conversion to a lower dimensional but more feature focused representation, a sparse but very long ranged context you could say.

I might be overstretching here, trying to see reason out arbitrary parallels that just happen to exist. There is definitely a long term memory, that’s automatically built out from things that I care about or interact with subconsciously but for the immediate task, like writing this note for example, I need a working short term memory and while I’ll retain the specific of this in some form or idea simply because the ideas in here have been brewing for quite some time, I’ll forget the specific structure and order of things in here, even if I remember them now as I’m working on it; so for the immediate task there is a short context window with a higher recall ability. What I’m tying to get at here is multi-tasking, which we are quite bad at naturally and can only improve at it somewhat via organisation improvements — a cost of context switch. I think of this as a prefix caching we do in our recent thoughts ( I was initially going to quote the model loading and switching argument but that’s obviously wrong ), the chain of thought being built is recent. If you keep on quickly switching to entirely different tasks ( “prefixes” ), you need to first recall what you were doing, and then reason what was next ( aka do the prefix building ) before you can actually make progress on the task ( emit next tokens ). Now of course people can easily juggle multiple tasks, so can LLMs, the same prefix ( “chat” ) can have several different things going on in it, but there is a clear performance decline; we can see the same for ourselves as well. I initially thought this fast switching would be a skill I can learn, and tried a couple things around it, but I now believe that whatever I’m doing is improving the “harness” which can only take me so far. There is no substitute to focused high signal work on a single task for hours on end — because in the end you are able to use longer and longer context dedicated to a single task, there is no mental drain ( confusion ) on the prefix re-writing or extracting bits for “this” task again and again, and in the end you are able to create a higher, much more dense long term representation of whatever it is that you were trying to learn.

This “confusion” aspect from trying to do multiple things at once, I want to elaborate on. Even if all I’m doing is just reviewing code and steering LLMs on different projects, which is simpler cognitively I would say; I get a mental fatigure and tiredness very very fast from this, the work done is also not the best that could’ve been, and there is a greater tendency to get stuck on some low signal task taking away all the gains of this parallelism that you’re trying to do.

I think that was all that I wasnted to write about today, I’ll keep on adding more stuff here as and when I go through the days but just having something like this written, I hope, would ensure I do not make the same mistakes again and again. For example, there have been weeks where I have an idea that in that week or a couple few folowing it, is the self evident truth of nature for me, but a few more weeks and I won’t even be able to recall it’s existence, just that maybe “something” has been lost from me, only to re-discover it a couple months from now. This cycle of keeping onto the same mistakes and improvements needs to stop; and by having written this note, I’ll never “re-discover” the same things about me.

In fact, you know, this very act of writing this note, to distill key learnings is what is also a crucial component of most agent harnesses, so that they do not end up rediscovering the same things again and again. I feel like all these things were always out in the open, there’s nothing groundbreaking about it, but it’s just that I’m able to appreciate their significance more and more now that I can see them working out nicely in an external system. We learn by osmosis after all; maybe I would’ve been able to realise all this sooner had I been in company of better humans — but well, that’s not needed anymore I guess.