Why do members here think DL-based methods can't achieve AGI?

If you mean like, the sensory input is ordering the command, I agree except I’m unsure about the semantics. The right semantics depend on the context and implications, which I don’t know.

Perceptual detection doesn’t necessarily originate in the thalamus.

L5tt cells burst to trigger perceptual detection [1]. Detection is highly impaired if their synapses are silenced in thalamus or superior colliculus, and also impaired for striatum.

Most L5tt cells can burst upon a stimulus, with no task or tangible reward [2]. Reward could be involved, but it doesn’t make a critical difference, so it would probably work without reward.

Now, if you just talking about motor commands, that’s much more iffy. They do burst like that even if they project to certain parts of the brainstem [2], but that study didn’t check their motor output [3].

I absolutely agree in spirit. I imagine it’d still do stuff with all its sensory inputs gone, like predict noise.

The evidence for this is scattered through many papers. It would be handy if this was the center of study in a paper but alas, this is usually mentioned in passing when showering attention on the prefrontal cortex, and the ones that are all about the basal ganglia generally don’t pay much attention to what is going on in the prefrontal cortex.

Sigh.

I have dozens of papers related to this area of study but these few should give you an idea of why I think that there are loops of control passing from cortical to subcortical structures and back again. A careful reading biases the subcortical structures as being the initiator of these loops.

Human Volition - towards a neuroscience of will - Patrick Haggard

A key passage is:
In practice, preparatory brain activity may begin as early as researchers are able to look for it. For example, recent attempts to decode free choices using new algorithms suggest that neural preparation begins much earlier than was previously thought. Second, the preparatory activity of the preSMA must itself be caused. The brain’s circuits for voluntary action might consist of loops rather than linear chains that run back to an unspecified and uncaused cause (the ‘will’). Indeed, the input from the basal ganglia to the preSMA is thought to play a major part in the initiation of action. For example, patients with Parkinson’s disease, in whom the output from the basal ganglia to the preSMA is reduced, show less frequent and slower actions than healthy controls. Moreover, signals that predict a forthcoming voluntary response can be recorded some 2 s before movement onset from electrodes implanted in the basal ganglia — these signals thus precede the typical onset time of readiness potentials. The subcortical loop through the basal ganglia integrates a wide range of cortical signals to drive currently appropriate actions, whereas dopaminergic inputs from the substantia nigra to the striatum provide the possibility to modulate this drive according to patterns of reward. From this view, voluntary action is better characterized as a flexible and intelligent interaction with the animal’s current and historical context than as an uncaused initiation of action. The basal ganglia−preSMA circuit has a key role in this process.

Interactions among the medial prefrontal cortex, hippocampus and midline thalamus in emotional and cognitive processing in the rat

Neural Correlates for Apathy: Frontal-Prefrontal and Parietal Cortical- Subcortical Circuits
What happens when the subcortical command signals are not very strong?
I don’t care.

How Basal Ganglia Outputs Generate Behavior

Parallel basal ganglia circuits for voluntary and automatic behaviour to reach rewards

Goal-directed and habitual control in the basal ganglia
Note that there are multiple control pathways from the subcortex. This paper describes two control systems.
Progressive loss of the ascending dopaminergic projection in the basal ganglia is a fundamental pathological feature of Parkinson’s disease. Studies in animals and humans have identified spatially segregated functional territories in the basal ganglia for the control of goal-directed and habitual actions. In patients with Parkinson’s disease the loss of dopamine is predominantly in the posterior putamen, a region of the basal ganglia associated with the control of habitual behaviour. These patients may therefore be forced into a progressive reliance on the goal-directed mode of action control that is mediated by comparatively preserved processing in the rostromedial striatum.

Inhibitory Control of Prefrontal Cortex by the Claustrum

Why is this important? Subcortex regulates cortical activity in many ways.

More is less: a disinhibited prefrontal cortex impairs cognitive flexibility

I maintain that a big chunk of what the cortex does is explain the world in a way that makes sense to the dumb boss, and take the will of the dumb boss and elaborate that into more complex and useful behaviors.

4 Likes

I think I was wrong about the anchoring thing. It stemmed from a study measuring bursts but kinda not single spikes (calcium imaging). The location selectivity doesn’t actually require bursting.

So I no longer disagree that the cortex is passive.

2 Likes

It doesn’t matter what the original initiator is, the “control” adds influences accumulated along the way. And it modifies all prior influences / “initiators” with feedback, that’s value drift in conditioning. Go through enough loops and it’s likely to extinguish original initiators altogether.

Subcortical primitives may start as a boss, but then “instrumental” values take over. In terms of evolution, it’s all instrumental anyway :).

3 Likes

So my Chinese Room has two books? the active and passive ones.
The singled book interpretation is naively wrong: the book contains written responses for whatever messages might arrive from senses. Not even bugs are that simple.

In the dual book interpretation, the “active” book contains rules about both

  • how to update the “passive” book content based on its current content and recently arrived messages
  • how to generate output messages based on passive book’s recently updated content

Well… a kind of.

2 Likes

Yes and no.

In traditional networks the nodes are all of the same types of neurons.
If you look at networks like autoencoders, you see a “choke point” in the middle.
This is the place where the very different processing methods of the subcortex is located. To complete this proposed model I would extend the autoencoder to feed a substantial portion of the output back to the input for a partially closed loop. The senses complete the input portion, and motor drives complete the output. This configuration combines the strongest features of both types of processors.
Cortex plus subcortex

While cortex is very powerful it could end up spending forever digging around in higher dimensional space for the most trivial of problems. The “simple minded” subcortex keeps the cortex grounded and focused on the most important task at hand.

If you don’t explicitly build the subcortex into your AI you will have to emulate the functions in some other way. Something will have to prioritize multiple competing goals, establish a primary goal, task switch if necessary, seed the machine state, start the processing, monitor progress, recognize when a solution is “good enough” and initiate action when the proposed solution meets some internally defined goal. Rinse and repeat.

Evolution has settled on this arrangement of cortex/subcortex and tuned it over the long haul. Subcortex worked pretty well without cortex. It works even better with cortex. Before you reject this arrangement out of hand it may make sense to understand how it works and what it is doing.

2 Likes

Not really.
There is still the books, cortex, and the agent that runs the rules of the room, subcortex.
For some very basic input messages, the agent responds directly without bothering with the books.

1 Like

Thanks, those are quite useful. Frankly, I feel your pain. Note how most of the papers address Parkinson’s disease. The Hindawi article was particularly provocative–loved the rat fitted with an electronics package, brought the Borg to mind.

The passive/active debate with the cortex is completely irrelevant for me. A few decades ago one of my grad students got interested in ‘computing in memory’ and there was a brief flurry of excitement in that regard. That is exactly what the cortex is, a computing memory. Just fascinating.

2 Likes

I think you are talking about supervised learning, it has nothing to do with autoencoders.
As for the rest, we’ve been through this a bunch of times, so it’s probably hopeless.

Regarding the recurrent autoencoder picture you have, I think there is a deep problem with that. It’s more of a hunch.
With that middle layer or “representational embedding” if you want. I’ll call it simply “representation”. By definition It’s a set of features that can be used to reconstruct the input (within a desired accuracy)
But there is no rule stating autoencoders representation vector have to be narrow. There are autoencoders that fan out their representation.
They are not popular since in DL the larger the vectors on each layer the more computing power is required for each dot product step.
And there-s the implicit assumption that the shorter the representation vector the more it embeds “higher order features”.
Which aligns with the observation that “conscience layer” is apparently very narrow - the numbers vary studies say we can not be aware of more than a few things at the same time. This “few” varies from one to 3, 5 maybe 7 depending on who are we asking.

The problems with that model start when we want it to learn more and more things. The representation optimized for certain input domain(s), becomes too narrow. Is not sufficient to add extra points in the middle representation layer since each change affects all following decoder nodes. Then you have to retrain the whole model, and cannot reuse previously known representations. “Consciousness” (whatever that means) will not recognize new representation. And another problem is the representation vector size is nowhere near 7, 5 or 3.


That’s why my hunch is the autoencoder fans out in an extremely large, yet sparse representation space.
How sparse? Forget 5% or 1%. It squeezes the shit out of it until 1, 3 or 5 feature points remain active and THAT is what consciousness layer sees. “Oh a TV!”. “That’s Tom!”

TLDR the autoencoder’s representation layer might be as large as many columns in the brain, and all learning machine purpose is continuous sparsification up to the point the input can no more be “explained” (== acceptably reconstructed) by the few remaining active points.

1 Like

Does it actually fan out in a different temporal dimension, in a way which none of the current approaches use… after all are autoencoders really just temporal compression and dilation mechanisms ? They convert one temporal dimension into another and then back again ?

The feedback loop would be heavily “what if I do this” type information, sort of a decaying predictive vector of all possible actions.

Well, quite possible, as long as we don’t have an actual implementation to see that it works all plausible hypotheses are welcomed.
Or should be welcomed. The state of affairs nowadays is when someone reports some important improvement in a certain, narrow direction, everyone else rushes in “We made a bigger transformer. one more % over SOTA! We nailed it!”

1 Like

I disagree. Just because Google Assitant or Siri isn’t capable of holding complex conversations doesn’t mean other models (mostly DL) cannot too. They are not “brittle detection” models, but I suppose we can never really comment on what they are seeing the pace of interpretability.

To hold complex conversations, I think they’d have to understand the world. They’d need human-level general intelligence. At that point, they’ll be useful for a lot more than conversations.

2 Likes

Again, as someone who actually works in implementing solutions, the biggest detriment to me and others who are trying to apply these (very useful, though absolutely brittle) solutions in production is that so many maybe well meaning, maybe hype-jacking, and maybe profiteering people are misrepresenting everything that DL can do, the ease with which it can be accomplished, and making ill-formed blanket statements that all we need is “more data” without stopping to consider all the potentially flawed and biased base assumptions that data brings.

The data might be utter garbage, the variables completely unrelated to each other (or just happenstance correlations), mislabeled (if labeled at all), or might shift concept or use randomly throughout the dataset (where some developer kept changing their mind about what a column was supposed to be doing, its categorical vs. numerical nature, range, interval, etc.). And that’s just the data aspect of it. Then there’s the algorithms themselves that we use, which again, are just clever mathematical tricks to attempt to force a certain “shapes” or “boundaries” into the jumbled mess, where the algorithms and parameters we set, by nature of their numerical embodiment, have unintended effects on the output shape such as creating clusters and divisions between groups that really shouldn’t be there, and yet we accept it because 80% recall is “good enough” for a certain application.

A production Deep Learning system is (oversimplified) just a numerical manipulation through a set of fixed-weight matrices which feed into functions. Our ability to get these systems to train, even with “clean” data assumes that any real or actual relationship exists between the input variables. Having to update weights through the bruteforce of backpropagation, though it sometimes works, isn’t guaranteed to find a working, or even a good solution consistently. There’s a lot of randomness and non-deterministic behavior so that even with the same architecture, same shapes, same data, even same learning rate and other hyperparameters, you still might not consistently arrive at a working model which means that you’ll still spend more time, energy, electricity, all of it, just to attempt to maybe get a working model.

So it’s often fine when you can get it to work and make sure to buttress it with all the required constraints and expectations, but way too many folks and companies out there are making way too big of claims about the ability of DL to solve problems, much less lead to AGI. Often, the people who talk the loudest about it know the least about how to actually implement any of it. They’re just hucksters looking to make a profit off the hopes/dreams of the gullible and ignorant-but-well-meaning.

DL is applied calculus and it IS pretty neat when it works. But network-wide back-propagation is a terribly inefficient way to conduct learning which produces fascinating and still brittle results, and I’ll stick by that. Even those impressive massive models which have memorized troves of written data (GPT-3, for example) are still terribly brittle and temperamental beasts around which folks are working hard to place hand-written filters and limiters so that only semi-correct answers are allowed to fall out.

Simply repeating the flawed approach over and over again while scaling up to powerplant-dependent levels of electricity is not going to cut it. Instead we’ll need Numenta (and others) who are pushing the boundaries on biologically-mimicking systems with more efficient basic operations, a different approach to the math, corresponding ASICs, and a rethink of how we’re picking/choosing what connections to update and when.

Attention mechanisms help in DL, but if we take a step back, the entire HTM approach was already a multilayered, multi-headed attention mechanism long before the DL community ever considered attention heads.

4 Likes

I suppose we have different ballparks of what a complex conversation implies. IMHO models like LaMDA are usually decent enough to hold what I consider pretty complicated conversations. For some reason, everyone seems to hold GPT3 to be king of hill ignoring other models that exist :thinking:

Says who? I am not particularly holding a position here, but I would need actual citations for some of the energized arguments you’ve made. The aim is not to convince but to understand things in an unbiased lens.

I don’t really get the attitude of most people here. Is this some sort of a war, between the DL community and Numenta, the latter being the protagonist, the underdog who has persevered through difficulties and would free people of the dark forces of DL that has usurped the minds of scientists? :joy:

It is not particularly towards @MaxLee but a general assessment of the responses I read here. There seems to be a certain antagonism for other methods (which I repeat, is not held by everyone - some frequent posters especially stand out in their openness to new ideas).

There isn’t any major disagreement about the achievements of Numenta, merely that they have still failed to overturn any LeaderBoards or provide shining scores. Some papers have displayed pretty decent results against baselines which I am very happy to see, but the point that drags it down is that Numenta was founded about 20 years ago - and the rate at which DL has produced results is unmatched.

There is a certain hope that Numenta’s work was always a long-term investment, but really I suppose the GOFAI vibes ensures that results take priority over almost everything. It is perhaps not the most effective way to reach AGI, but it is a tried & tested methodology that has empowered other scientific fields for decades now.

1 Like

I dunno… if Geoffry Hinton, a guy who’s lived in this most of his life, thinks that backprop is flawed, I’ll weight his opinion on the matter highly. As for why others who’ve seen it don’t get it yet, well… most humans love their niche. Our brains, biologically, want to use the least amount of calories to accomplish their goal (staying alive, making/keeping relationships, earning money, entertainment, etc.). They may be brilliant experts in their fields, and there’s nothing wrong with that. But, we also have oddballs who aren’t satisfied and want to know everything about everything, even those seemingly impractical things that don’t serve any immediate purposes. Lucky for us, those folks exist and keep concatenating knowledge and experience to synthesize new ideas in art, mathematics, science, technology, etc., and keep us moving forward in our understanding of the universe. I’d politely say Hawkins is one of those oddballs. :slight_smile:

Without them saying specifically how much of their process required hyperparemter optimization, how many restarts (false starts, failed convergences, random experimentation), and other general activities that take place when trying to train a DL model, especially a massive transformer model, I feel fairly confident that these estimates are definitely the lower-bound of energy requirements. No way they simply had a straight shot and taught the model without issues.

https://www.reddit.com/r/MachineLearning/comments/htxjoq/d_gpt3_175b_energy_usage_estimate/

I wouldn’t describe it as “antagonism” per-se, but more a cleared-eyed acceptance or understanding of what these approaches can or can’t do, with a mild disgust towards the over-hyping of DL by profiteering self promoters whose only goal frequently seems to be fast unicorn funding raises before skipping out leaving others holding the bag. I have a love for bayesian, markov, spiking, tree-based, boosting, genetic, symbolic, etc… and the whole family of algorithms under the umbrella of AI. But this topic is specifically asking why it’s believed that DL won’t result in AGI, and I’ve put forth my thoughts, quite loudly. :slight_smile:

Speaking of DL and Hinton… they both just languished for ~20-30 years before data and hardware caught up. Seeing what HTM has been able to do with FPGAs (limited in gates and overall clock speed) already quite impresses me, as have their early experiments in applying sparsity to deep learning. Even if Numenta were to evaporate today, those seeds aren’t lost with them. I’ve a suspicion we’re about to see them bloom regardless, and taking inspiration and modelling from the brain was the source of it all.

1 Like

Here’s something that I think highlights a key tension in this discussion:

When I mention to folks in the DL world, what Numenta is working on, the overwhelming response is “But does it work? *Show me the results!” The field is deeply pragmatic, and willing to go along with approaches even if they are not elegant or efficient, simply because they actually accomplish the task.

In order to take neuro-AI work as anything more than idle tinkering, then, a group like Numenta needs to show that the methods they develop are on a path towards performing comparably at the tasks that DL methods have made progress on (such as text generation). The experiments with sparse NNs on FPGAs are a small step in that direction, although I would caution that comparisons between methods that rely on hardware-specific optimizations (ones that simply don’t work on the actual hardware we have, i.e. GPUs) and ones that are generally applicable are looked on with great skepticism.

That being said, I look forward to the day that non-backprop NN approaches start working as well as backprop NNs do today. That will be a grand occasion.

1 Like

that still doesn’t guarantee that future methods would have similar energy requirements. I agree with the points raised in the Reddit thread - ultimately, I believe any expense would have been worth it.

I hope so too :slight_smile:

I suppose they sharded considerably for hyperparameter-optimization and test runs anyways, so the usage wouldn’t be on-par the final 175B run.
Not that it matters much anyways…

Neuroscience is a long-term investment.


Possible-Probable, my black hen.
She lays eggs in the relative when.
She doesn’t lay eggs in the positive now
because she simply can’t postulate how.


If you define AGI as something that can replace a human in an industrial/factory setting, then GOFAI seems like a good approach. But if you define AGI as something that can act like a human, then you should study neuroscience.

3 Likes