Sometime last week I “took” a sticky notes pack from work, my intention then being to write my own bits and pieces of notes on it. It’s more than a week and I haven’t done any of that yet but a curious thought still lingers and I want to pen it down today.
To do long form reasoning, LLMs started using very information dense languages like Chinese or a baser form of English with all grammar thrown out the window. I can understand the idea somewhat, if you were to define a region in the high dimensional space they operate in, then just a few words within that region would allow it to guide, most words are not attended to. Once you start getting into lossy territory, all you would need is just a bag ( not really for I still assume a rough order to it ) of words.
Before I finish that line of thinking, I need to pen down another. For people who themselves are capable and have put in the time and effort, not very much is needed for them to perform significantly better. Take the case of some chess grandmaster, they won’t fall for the 99% of the cases, having put in the time and effort to ensure that. All they need to be the best in the world is just a small hint, even a warning that a certain next move is very important without revealing the move itself is enough to tip the balance. It’s the same with memory and day to day life, it’s fractals all the way.
When I forget something, I never forget it in whole, I always have some vague recollection of it and just “one glance” is enough to make me recount it. Sometime forget about say a topic x that I spend time reading up on, only to encounter a trigger word which then lets me recount what I read earlier, atleast something of it. I want to call it the information dense representation of an idea. For a topic, if you can just jot it down in an as condensed and lossy way as possible, such that the most important of word triggers survive, effectively making the topic nothing but a bag of words, then it won’t mean anything to anyone BUT you. To you one glance at those words is enough. Neurons I guess work the same way, once you are close to a region, each such “trigger” word defines the path, the brain can “walk” a couple steps along that path and then you once again give it another trigger word ( a direction to another neuron, same as the coordinate in the high dimensional space ), with enough such words and some practice, it’ll be enough to have “notes” that are so so so compressed, they are at once both gibberish and yet everything; for now we are not trying to store the information on paper, such that reading it in full will recount you the details. You have to have read it in full and have the details in your mind as a prior, the “word bag” then only serves as “query” words to fetch the information from your brain itself.