This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:
gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200
MENENIUS:
'Though all at once canq
MARCIUS:
Pray now, nocamest thou to a morsel.
LARTIUS:
Hence, and
I' the end admire, where G
again; and after it ag .
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.
The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.
sourdecor 2 hours ago [-]
I don't know how this is related, but it reminds me of how I have always believed compression to be the ultimate sign of intelligence. If you can reduce something while keeping comprehension, you are finding more abstract symbols to represent the information of the original source.
How can things be compressed without losing information or structure?
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
doctoboggan 4 hours ago [-]
Other than what the others mentioned about finding more efficient representations, you can also compress by pre-agreeing on some common terminology.
In many ways, we are communicating using a compressed channel (words) since we both have pre-agreed on the meaning of these words.
gchamonlive 3 hours ago [-]
That and also that agreements can evolve with time and also within a discussion, so a basic level of agreement is necessary, but complete consensus about the meaning of all words is unnecessary and often unproductive to communications held in good faith.
nkmnz 8 hours ago [-]
You analyze the frequency of combinations of bytes, then replace those with high frequency with pointers to a single instance.
selicos 4 hours ago [-]
Deduplication ^
mpalmer 8 hours ago [-]
How can things be compressed without losing information or structure?
Because the initial content is rarely the most efficient representation, so it's possible to store fewer bytes that can deterministically be converted into the original.
Like for text, what would that involve?
Most compression algos don't care what information you're compressing. All they see (all they need to see) is bytes. It ends up being way more sophisticated than removing repeated periods and whitespace.
Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
gchamonlive 3 hours ago [-]
The initial content is the most efficient representation but not for persistent digital storage and that's an important distinction because what we call sparse or inneficient is actually quite efficient, but only if you consider human consumption as the optimization target.
red75prime 15 hours ago [-]
It reminds me of "At the time we drew boxes labeled 'perception', 'cognition' with arrows between them." An imprecise quote that I can't place.
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
bananaflag 15 hours ago [-]
You mean "Artificial Intelligence meets Natural Stupidity" by Drew McDermott
McDermott's A Critique of Pure Reason pretty much captured all of the misgivings I had about "Good Old-Fashioned Artificial Intelligence", which was slightly unfortunate as I was trying to complete a PhD in that very area at the time (around 1990 or so...)
bananaflag 12 hours ago [-]
Did you complete it?
arethuza 8 hours ago [-]
No - there were lots of other interesting things going on in the early 90s that I managed to pivot to. Never regretted not completing as I had given up any desire to work in academia by that point.
Most humans have weak meta-cognition, a large percentage doesn't have verbal thoughts.
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
arbirk 10 hours ago [-]
Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering
creativeSlumber 14 hours ago [-]
How relevant is this fast/slow thinking thing with regards to current frontier models?
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
Shorel 12 hours ago [-]
That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language.
It can be said that language is a tool for the serialization (writing) and deserialization (reading) of human ideas. It is also an incredible useful and powerful tool by itself.
This last sentence has been proved true by LLMs themselves.
However, since it is working on the serialized version of ideas, I agree with you in that's not the optimal way to think and something not serialized (maybe world models) can be invented that's better for thinking.
All this in no way diminishes the usefulness of language and of automated language generation.
virgilp 10 hours ago [-]
> That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
Shorel 7 hours ago [-]
I can't agree about the "unable to recognize and process it", simply because that idea is totally contrary to my own experience.
I have in fact many memories which have emotions in them, without words or other external elements.
However, seeing that language serialization seems to enable a vastly extended memory (entire sagas remembered as songs), it is understandable that something is gained by serialization of emotional experiences, just as something more immediate is lost.
globalise83 7 hours ago [-]
Sounds uncannily similar to the pseudofacts you hear a lot in Neurolinguistic Programming training courses for sales reps. Would ask for a scientific publication reference on that one.
Shorel 7 hours ago [-]
It seems to me that idea is rooted in social consensus.
Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.
Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.
I will not comment on who can belong to each group, that's an exercise for the reader.
theptip 4 hours ago [-]
Very relevant. Modern models use CoT to do “slow thinking” and this enables them to achieve much greater performance. You can also turn off thinking and answer directly which is quite similar to “fast thinking”, good at approximate maths, not capable of algorithms, etc.
Of course the shapes of what an AI can do in fast vs slow are quite different.
glial 6 hours ago [-]
'fast' means executing a policy, that is, a state-action mapping. A trained RL model does this.
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
creativeSlumber 5 hours ago [-]
Wouldn't you need a classifier to even decide if it is system 1 or 2? How capable does this classier need to be?
glial 5 hours ago [-]
I think of System 1 as a hash map. If you have a map, and see a new state/key whose action/value is not defined in the map, you have to go with System 2.
jhrmnn 10 hours ago [-]
No clue what’s the consensus on this but my internal mental model is absolutely that LLM AI is pure fast mode, no slow mode. The “reasoning” loops are an attempt to mimic the slow mode but ultimately it doesn’t really work. I’m curious about the recent maths advances though, they seem to possibly challenge this.
hoppp 6 hours ago [-]
Its just hype talk. Slow mode is conscious in humans, fast is subconscious, so we need to discuss AI consciousness to talk about system 2 thinking.
So the entire debate is fubar.
Retric 13 hours ago [-]
You can ask a model for output directly and stop, or you can recursively ask it to keep refining the output.
That seems to fit the fast vs slow model of human thought reasonably well.
schainks 3 hours ago [-]
Not quite. The better analogy for you is the "thinking" setting on your model.
usernametaken29 12 hours ago [-]
> You can ask a model for output directly and stop
That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here
Retric 12 hours ago [-]
That’s a function of the amount of processing power involved not the underlying architecture of decision making.
usernametaken29 10 hours ago [-]
In terms of making an LLM faster but not in terms of meta-cognition. System 1 thinking as defined by Kahneman doesn’t have 100000x more compute than System 2, it is actually the opposite. That completely contradicts your claim
pixl97 6 hours ago [-]
>System 1 thinking as defined by Kahneman doesn’t have 100000x more compute than System 2, it is actually the opposite.
When are you measuring?
Systems 1 thinking is closer to precomputed tables in some ways. That is by evolution or massive amounts of training your neural network has a narrow fast path it can execute with as little compute at execution as needed.
Retric 4 hours ago [-]
Slower paths means loops here for humans where the output of a neuron gets feed back into itself. The fastest path = a feed forward neural network without loops.
The ratio between a single pass and multiple passes is unchanged when you throw more processing power at both.
In a human 30ms vs 3-4 seconds is a 1:100 ratio. Single vs multiple passes with an LLM varies but a 1:100 ratio isn’t unrealistic. So with enough compute and the right workload single vs multiple pass LLM could sit in that exact same 30ms vs 3-4 second timeframe.
hoppp 7 hours ago [-]
Multiple passes doesn't make it system 2.
The defining characteristic of system 2 is consciousness which is expensive and slows down the system.
So to talk about system 2 in AI we need to talk about consciousness. As long as AI is not officially conscious there is no System 2 thinking implemented
Retric 4 hours ago [-]
Consciousness is ill defined hogwash when used in such descriptions.
The process of internal refinement without external action however fits.
hoppp 3 hours ago [-]
Using system 1 and 2 for AI is ill defined hogwash indeed.
The systems apply for humans and describe conscious and subconscious processing. Without that the entire reference to thinking fast and slow is bullshit.
Retric 2 hours ago [-]
System 1 and 2 is a description of physical process in people.
Using conscious vs subconscious processing is not however a meaningful definition. Subconscious processing isn’t necessarily fast. Visual processing and sensory integration can be quite slow without any conscious input.
ghm2199 11 hours ago [-]
Structurally speaking we learn nothing like AI, we don't use vast amounts of information to pick up completely new skills. We also make decisions by using prior knowledge and emotions.The latter part is important, Thinking fast and slow cannot operate in a world of AIs as they stand today unless we are willing to grant them rights — because you have to teach them to make decisions based on all kinds of emotions — which is tricky at best.
TeMPOraL 11 hours ago [-]
> we don't use vast amounts of information to pick up completely new skills.
Except, we do.
hoppp 7 hours ago [-]
Not relevant at all. Its a way to hype.
emp17344 11 hours ago [-]
Seems like a terrible idea in the first place to build an entire organization around a single pop-sci book, but that’s just me.
glial 6 hours ago [-]
System 1/2 is the pop version, but the fast/slow distinction is prevalent in both RL and computational cognitive science, sometimes going under different names: procedural/deliberative, model-free/model-based, automatic/controlled, associative/rule-based, autonomous/algorithmic, etc.
TeMPOraL 11 hours ago [-]
It's indeed a terrible idea - as in, it's great. You get benefits of cross-marketing: you ride on a popularity of a well-known book, and as you also drive more sales of it, even if you don't have a deal and don't benefit from that directly, you strengthen the loop and solidify your brand.
This choice doesn't really constrain what the organization can do, either. Pop-sci books have plenty of wiggle room in interpretation, and afford a lot of "you're holding it wrong" dismissals of criticism, that with a bit of clever copywriting, the organization can do absolutely anything and still claim it's embodying the framework/theory of the book.
olgava 13 hours ago [-]
[dead]
anotha_one 12 hours ago [-]
[dead]
crorella 16 hours ago [-]
It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.
alansaber 11 hours ago [-]
Given how LLMs access compressed knowledge from their model weights, the similarities make sense
mdk2578 13 hours ago [-]
[flagged]
zfoong 14 hours ago [-]
At least this is written before ChatGPT.
vist_orn 13 hours ago [-]
Trying to get LLMs to 'think about their thinking' is my daily struggle. This paper nails why it's so critical.
readthenotes1 12 hours ago [-]
If I recall correctly, all that fast and slow business has been debunked as yet more non-replicable pop psychology.
I shouldn't be surprised that it shows up in a screed on AI
bbor 16 hours ago [-]
This is still a great paper, but it's missing the second axis of the quadric -- if the only two options are thinking fast or thinking about thinking, that leaves no room for thinking slow yet deliberately, AKA selfconsciousness. See https://www.gutenberg.org/cache/epub/4280/pg4280-images.html for details
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
BOOSTERHIDROGEN 11 hours ago [-]
Isn't system 2 by itself semantically means thinking by thinking
jannyfer 17 hours ago [-]
> submitted Oct 5 2021
(In case people miss that before discussion)
tomhow 15 hours ago [-]
Updated, thanks!
locitra 5 hours ago [-]
[flagged]
ot4t 7 hours ago [-]
[dead]
aidiscoverywire 15 hours ago [-]
[flagged]
BigDogAU2026 10 hours ago [-]
[flagged]
tug2024 6 hours ago [-]
[dead]
simianwords 15 hours ago [-]
This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?
edit: why is this downvoted?
globnomulous 13 hours ago [-]
It's being downvoted, I think, for a few reasons:
* The person who posted it likely posted it not as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on what you're calling its out-of-dateness.
* You say "this has been solved" without defining what "this" is.
* Your description of the solution -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model and past experience. "Effort" isn't so much the issue as types of effort using different systems, modeled specifically after Kahneman's idea of fast and slow thinking.
* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.
In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.
simianwords 13 hours ago [-]
> it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model.
do you even know how adaptive reasoning works?
globnomulous 13 hours ago [-]
Godspeed to you in your efforts to contribute productively to Hacker News threads.
simianwords 13 hours ago [-]
> For the first time, GPT‑5.1 Instant can use adaptive reasoning to decide when to think before responding to more challenging questions, resulting in more thorough and accurate answers, while still responding quickly. This is reflected in significant improvements on math and coding evaluations like AIME 2025 and Codeforces.
It says literally the thing you wanted from system 2. Its almost exactly that.
This is what you said btw:
"it's that the system itself decides how to reason based on the nature of the problem it faces"
globnomulous 3 hours ago [-]
My second response was sassier than it should have been. Sorry for that. I'm not making any claims about whether or not GPT implemented this. My point was narrower in a way that is probably annoying: you asked why you were being downvoted; I tried to explain that, as far as I could tell, the adaptive system you explicitly described (regardless of what adaptive reasoning actually is), where humans choose the reasoning level, isn't the one that the paper proposes -- and that this may be one of the reasons for the downvotes.
I don't know whether that's why anybody downvoted you, but I think that's probably part of the reason. On the other hand, the Hacker News guidelines ask us to respond to the strongest interpretation of a comment or post. In that spirit, just as you could have said something to the effect of "Yep, AI researchers learned from behavioral economics and research into human decision making and here's where OpenAI mentions it," I could have said just "the relevance may not be that it's a new idea but rather that it's a neat idea, to the person posting it at least. Maybe someone who knows AI systems more deeply wouldn't find it interesting. I do though!" That's a bit of hypocrisy on my part, I think.
Anyhow, for the record, yep, I knew of adaptive reasoning, but, no, I didn't know that it matched what this paper proposes so closely. Thanks for the correction!
lelanthran 13 hours ago [-]
> This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?
Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.
simianwords 13 hours ago [-]
Asking earnestly, I don’t know what you mean by this reply. I know what system 1 and 2 is. But this has already been solved using same model.
lelanthran 13 hours ago [-]
> I know what system 1 and 2 is. But this has already been solved using same model.
No, it hasn't. Maybe you have a different definition of System 1 and System 2. I last read the book well over a decade ago (2011, maybe? 2012?), but System 1 and System 2 are different systems. IOW, System 2 is not a more computational version of System 1.
The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
In computery terms, System 1 runs in O(1) time, System 2 runs in O(log n) (or maybe just O(n)) time.
This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer"). We don't have LLMs that do that. We have System 2 - run in O(log n) time and produce an answer.
System 1 is completely bereft of thought.
hoppp 6 hours ago [-]
The difference in humans is consciousness. System 1 is subconscious and fast, system 2 is conscious and slow.
System 1 can do multiple passes and all that, the speed is related to how aware are you of the computation occuring and system 2 must be consciously managed.
Example: You write a quick reply to a comment - thats system 1
You add two four digit numbers in your head - that will be system 2 unless you are very good, then it can be system 1 also
Why? Because memory allocations and addition computations are manually managed.
creativeSlumber 5 hours ago [-]
is this just a fancy way of saying that sometimes humans makes quick heuristic based decisions, and sometimes they think through things thoroughly?
I don't think there's any problem category that is strictly a quick heuristic decision or something that you will think through thoroughly always. I think it's more about how much time you have. If you don't have time you'll make a quick heuristic based decision. If you have more time you will think through it more. Imagine you're driving and suddenly like a branch falls on the road right in front of you. And you need to avoid it. Your brain will quickly use a heuristic based approach to avoid the branch. But on the other hand, if this branch was already there on the road and you saw it from far away, you will probably take a lot more time to think through and figure out which path you need to take.
pixl97 6 hours ago [-]
> but System 1 and System 2 are different systems.
They also share a lot of overlap in brain structures and they interplay while executing. A thought can start out Sys1 and quickly migrate to Sys2 as the pattern fails to match. Or a System 2 chain of thinking can be made from a bunch of smaller system 1 actions. Heck, in the middle of a system 2 thought you can plunge into system 1 system/actions. It's more of a who matches the pattern up with reality the fastest and acts on it.
> We don't have LLMs that do that. We have System 2
Eh. LLMs are system 1 thinkers by default. "Quick" response with no reflection is where LLMs started. It's later we added Chain of Thought and reflection layers and all kinds of other things like harnesses and agent training to make them act like system 2 thinkers. Of course we have other technologies being tested on LLMs these days like Dynamic Sparse Attention that likely match more of your thoughts on what system 1 thinking is too. Where the network doesn't have to parse the full context of the prompt and instead pattern matches with a much smaller percentage of the input giving responses back in ms versus seconds.
simianwords 13 hours ago [-]
> The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
No, system 2 is the emergent capability to reason and increase the space of places to find the answer. Forget the paper's proposal, and look at the problem it is trying to solve. Ability to give quick answers, ability to give thought out answers, and the ability to know when to choose what. Adaptive reasoning does all three.
> This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer").
No, I don't think we humans use o(1) to for understanding 1000 tokens or 2 tokens. I simply don't think that's the case. There's a new model called "Jev" and it is literally named System 1 (from the book) and even it is billed per input token.
bpshaver 15 hours ago [-]
Surely that is obvious
sinuhe69 15 hours ago [-]
2021. Please remember the rule of HN to add the year if it’s not actual.
It had this to say in the linked post:
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.
https://en.wikipedia.org/wiki/Hutter_Prize
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
In many ways, we are communicating using a compressed channel (words) since we both have pre-agreed on the meaning of these words.
Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
https://dl.acm.org/doi/pdf/10.1145/1045339.1045340
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.
Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.
I will not comment on who can belong to each group, that's an exercise for the reader.
Of course the shapes of what an AI can do in fast vs slow are quite different.
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
So the entire debate is fubar.
That seems to fit the fast vs slow model of human thought reasonably well.
That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here
When are you measuring?
Systems 1 thinking is closer to precomputed tables in some ways. That is by evolution or massive amounts of training your neural network has a narrow fast path it can execute with as little compute at execution as needed.
LLM’s operate strictly feed forward neural networks.
In a human 30ms vs 3-4 seconds is a 1:100 ratio. Single vs multiple passes with an LLM varies but a 1:100 ratio isn’t unrealistic. So with enough compute and the right workload single vs multiple pass LLM could sit in that exact same 30ms vs 3-4 second timeframe.
So to talk about system 2 in AI we need to talk about consciousness. As long as AI is not officially conscious there is no System 2 thinking implemented
The process of internal refinement without external action however fits.
The systems apply for humans and describe conscious and subconscious processing. Without that the entire reference to thinking fast and slow is bullshit.
Using conscious vs subconscious processing is not however a meaningful definition. Subconscious processing isn’t necessarily fast. Visual processing and sensory integration can be quite slow without any conscious input.
Except, we do.
This choice doesn't really constrain what the organization can do, either. Pop-sci books have plenty of wiggle room in interpretation, and afford a lot of "you're holding it wrong" dismissals of criticism, that with a bit of clever copywriting, the organization can do absolutely anything and still claim it's embodying the framework/theory of the book.
I shouldn't be surprised that it shows up in a screed on AI
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
(In case people miss that before discussion)
edit: why is this downvoted?
* The person who posted it likely posted it not as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on what you're calling its out-of-dateness.
* You say "this has been solved" without defining what "this" is.
* Your description of the solution -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model and past experience. "Effort" isn't so much the issue as types of effort using different systems, modeled specifically after Kahneman's idea of fast and slow thinking.
* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.
In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.
do you even know how adaptive reasoning works?
https://openai.com/index/gpt-5-1/
It says literally the thing you wanted from system 2. Its almost exactly that.
This is what you said btw:
"it's that the system itself decides how to reason based on the nature of the problem it faces"
I don't know whether that's why anybody downvoted you, but I think that's probably part of the reason. On the other hand, the Hacker News guidelines ask us to respond to the strongest interpretation of a comment or post. In that spirit, just as you could have said something to the effect of "Yep, AI researchers learned from behavioral economics and research into human decision making and here's where OpenAI mentions it," I could have said just "the relevance may not be that it's a new idea but rather that it's a neat idea, to the person posting it at least. Maybe someone who knows AI systems more deeply wouldn't find it interesting. I do though!" That's a bit of hypocrisy on my part, I think.
Anyhow, for the record, yep, I knew of adaptive reasoning, but, no, I didn't know that it matched what this paper proposes so closely. Thanks for the correction!
Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.
No, it hasn't. Maybe you have a different definition of System 1 and System 2. I last read the book well over a decade ago (2011, maybe? 2012?), but System 1 and System 2 are different systems. IOW, System 2 is not a more computational version of System 1.
The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
In computery terms, System 1 runs in O(1) time, System 2 runs in O(log n) (or maybe just O(n)) time.
This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer"). We don't have LLMs that do that. We have System 2 - run in O(log n) time and produce an answer.
System 1 is completely bereft of thought.
System 1 can do multiple passes and all that, the speed is related to how aware are you of the computation occuring and system 2 must be consciously managed.
Example: You write a quick reply to a comment - thats system 1
You add two four digit numbers in your head - that will be system 2 unless you are very good, then it can be system 1 also
Why? Because memory allocations and addition computations are manually managed.
I don't think there's any problem category that is strictly a quick heuristic decision or something that you will think through thoroughly always. I think it's more about how much time you have. If you don't have time you'll make a quick heuristic based decision. If you have more time you will think through it more. Imagine you're driving and suddenly like a branch falls on the road right in front of you. And you need to avoid it. Your brain will quickly use a heuristic based approach to avoid the branch. But on the other hand, if this branch was already there on the road and you saw it from far away, you will probably take a lot more time to think through and figure out which path you need to take.
They also share a lot of overlap in brain structures and they interplay while executing. A thought can start out Sys1 and quickly migrate to Sys2 as the pattern fails to match. Or a System 2 chain of thinking can be made from a bunch of smaller system 1 actions. Heck, in the middle of a system 2 thought you can plunge into system 1 system/actions. It's more of a who matches the pattern up with reality the fastest and acts on it.
> We don't have LLMs that do that. We have System 2
Eh. LLMs are system 1 thinkers by default. "Quick" response with no reflection is where LLMs started. It's later we added Chain of Thought and reflection layers and all kinds of other things like harnesses and agent training to make them act like system 2 thinkers. Of course we have other technologies being tested on LLMs these days like Dynamic Sparse Attention that likely match more of your thoughts on what system 1 thinking is too. Where the network doesn't have to parse the full context of the prompt and instead pattern matches with a much smaller percentage of the input giving responses back in ms versus seconds.
No, system 2 is the emergent capability to reason and increase the space of places to find the answer. Forget the paper's proposal, and look at the problem it is trying to solve. Ability to give quick answers, ability to give thought out answers, and the ability to know when to choose what. Adaptive reasoning does all three.
> This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer").
No, I don't think we humans use o(1) to for understanding 1000 tokens or 2 tokens. I simply don't think that's the case. There's a new model called "Jev" and it is literally named System 1 (from the book) and even it is billed per input token.