Progress in the Analysis of Cortical Circuit Implementations

Hello all!

FYI, this post is mirrored on my blog although I’m writing directly to you folks.

Some time ago I coined the term Discrete Cortical Circuits and described the distinguishing characteristics that many have in common. My original list of DCC characteristics included the following:

  • neuron activations are binary
  • synaptic weights are binary
  • winner-take-all activation strategy for neurons
  • input data must be encoded to binary

You can keep adding more DCC characteristics to the list:

  • Hebbian learning
  • high-dimensional (1000s of bits)
  • sparsity (< 20% of bits active)
  • structure in the binary codes (order, coordinates, topology, groups)

I’ve selected and been analyzing, what I consider, four different implementations of DCCs. None of these four satisfy all the characteristics above. However, the most interesting thing about these DCCs is in how they violate the criteria and what that changes for the performance of the algorithms. The four different DCCs I considered are: HTM, Sparsey, SPH, and BrainBlocks.

I had vague ideas about what i wanted to accomplish analyzing and comparing these architectures. I was hoping that I would somehow find the core commonalities or essential features that would enable a new fundamental understanding of intelligence. I’m not sure I’m there yet, but I certainly have discovered some knowledge that I hope to share with you.

One of the things I’ve been building is what I call the DCC Reference Model which is accompanied by a taxonomy of architectural decisions you can make. An example of what I mean by this would be, deciding how you are going to select which cells to activate and in what structure they can be broken down by:

  • grouping: partitioning your cells into groups or keeping them unpartitioned (one group)
  • winner-selection: a single winner per group, k winners per group, or at-most-k winners per group.

There are many other architectural facets, but I will come back to them later.

Another thing I wanted to do was build a common interactive sandbox so that people could play with these algorithms, easily assemble them, scale to extremely large networks, and visualize their operation. For that purpose, I’ve been working on my DCCcore framework which is something hasn’t been released yet. Not only does it have a high-performance Rust backend, but I have created a browser-based frontend. You can now run DCCs in your own browser with amazing speed and efficiency. They also provide real-time graphics and interactivity.

I initially announced my web-based dashboard project on my blog. Since then, I have been posting new iterations of my visualization and interactive applications on my DCC public repo. No source code for these yet, but the WebAssembly apps can be embedded in any web page or viewed directly. Feel free to check out some of my older web demos.

Anyway, as part of the effort for the comparison and analysis of other DCC architectures, I ended up porting the existing implementations from their C++ or Java languages to the Rust language. This not only lets me embed them into the browser with WebAssembly, but it also helped me understand the different DCC architectures better. I definitely had some mechanical help doing this with the aid of your friendly neighborhood LLM agent. One of the consequences of this is that the resultant product is heavily annotated and documented to ensure clear human and machine understanding of the code, with a set of standards and tests to ensure fidelity to the original project. You can dive into the doc subdirectories to look at the reframing of these projects in words other than the developers. I think even this latter part is very useful since developers often have a hard time explaining their own work since they’re so deep in it. I know I have this problem too!

I have released the source code and repos of the compared DCCs. Though, I was the co-developer of BrainBlocks, I haven’t released the ported Rust code since there are some IP issues I need to clear with the company I worked at when writing this. Other than this, I have used the licenses provided by the project or got written permission where the reference code was not publicly available.

Here are the three repos I have released:

The original reference implementation for HTM is C++ with Python bindings, for Sparsey is Java, and for SPH is C++. The Rust repos are fully stand-alone implementations with example code, documentation, and tests provided. You may encounter the occasional error or AI hallucination, but I’ve inspected the work myself and enforced a lot of quality on the end product.

I would welcome anyone who dives into these projects to offer any corrections or improvements you might have. I’d also welcome any further example code or implementation of features that I didn’t provide. The repos have to conform to a few minimal structural standards so that they can be fetched and built with my framework and web dashboard. Other than that, these projects stand on their own.

Hi, it seems there-s quite a lot of work there.

I see you mention DCC versions for HTM, Sparsey, SPH but what about brainblocks? Or DCC itself is the update to it?


I like you have RL examples for SPH, I think RL is the first approach one should consider for testing their AI frameworks. I see you have more examples than Ogma provides, or I didn’t search enough? I’ll have to look at them anyway,

Personally I did a lot of experiments with SPH and found

  1. it has much better high-order prediction results than HTM
  2. SPH is only one whose concept is directly aimed for RL.

What’s about your experience, @jacobeverist ?

@thanh-binh.to Yes, that is my experience too. However, I’ve been trying to do side-by-side implementations of the exact same type of RL applications and seeing where things are different. So far, I’ve SPH is the clear winner for RL applications, but I’m still trying to understand the problem space better. I can get my own algorithms to solve some of the problems but other problems I can’t solve at all.

One of the key advantages that SPH has over other DCC algorithms is that they still use continuous weights, which gives them a smooth gradient between the current output and the target output. This is not something that exists currently with binary weights. I’m working on solution for this that doesn’t sacrifice binary weights, and we’ll see how it goes.

@cezar_t As I mentioned above:

Here’s some food for thought based on the BrainStem architecture that might help with analyzing and refining Discrete Cortical Circuits.

First off, instead of relying on static Winner-Take-All thresholds or uniform Hebbian learning rules across all processing steps, BrainStem uses a system of digital neuromodulators like Dopamine, Serotonin, Glutamate, GABA, Noradrenaline, and Acetylcholine to continuously adapt network dynamics. These neuromodulatory levels actively shift the system between different operational regimes such as Exploration, Precision, and Protection or Consolidation. Key parameters like learning rate, error weighting, exploration pressure, revision pressure, and filter strictness are dynamically tuned based on feedback and neuromodulatory state rather than staying fixed.

Second, to prevent catastrophic forgetting and avoid burning in self generated errors or hallucinations, BrainStem introduces an offline sleep consolidation phase following active learning loops. During replay, the system uses an interleaved batching strategy that mixes novel candidate patterns with highly stable anchor patterns. It implements a slow wave sleep substructure that runs candidates through simulated up states and down states. Candidate activations undergo multi cycle reactivation where selection pressure, modulated by GABA and Glutamate gains, determines whether representations are reinforced, weakened, or pruned.

Third, to enforce high sparsity under 20 percent and prevent runaway activation in dense or binary networks, BrainStem continuously computes an Excitation Inhibition balance through the dynamic interplay of excitatory Glutamate drive and inhibitory GABA drive. When activation weights or parameters reach boundary limits over consecutive cycles, BrainStem applies sliding threshold homeostasis and meta metaplasticity to make counter directional adjustments easier. If parameter saturation co occurs with low effectiveness, the system executes a controlled slow wave downscaling or bias renormalization to pull saturated weights gently back toward mid range targets.

Finally, to ensure that new sparse binary codes or structured circuits represent genuine patterns rather than noise, candidate structures must undergo guarded hypothesis graduation. An unconfirmed representation cannot immediately be promoted to a permanent state and must survive a minimum number of slow wave sleep consolidation cycles. Proposed structural and plasticity changes must pass a critic snapshot gate that verifies parameter deviations stay within defined safety tolerances before permanent graduation is granted.

—-
AI Part

Here is a structured comparison report and actionable recommendations mapping the core requirements of Discrete Cortical Circuits (DCCs) directly to the architectural solutions developed in BrainStem.


Structured Comparison: DCC Taxonomy vs. BrainStem Architecture

DCC Core Criterion Standard DCC Challenge / Violation BrainStem Architectural Solution Grounding & Source Mechanism
Dynamic Learning & Plasticity Rigid, static Hebbian update rules and fixed Winner-Take-All (WTA) thresholds lead to catastrophic forgetting or premature convergence. Digital Neuromodulation & Regime Control: Dynamically scales learning rates, exploration pressure, inhibition levels, and consolidation gains based on global system states. Multi-agent neuromodulation (Dopamine, Serotonin, Glutamate, GABA, Noradrenaline, Acetylcholine, Adenosine, Endocannabinoids) shifts the system between Exploration, Precision, and Consolidation regimes .
High Sparsity (< 20%) & WTA Saturating activations or runaway excitation in high-dimensional binary representations can break sparsity limits. Adaptive E/I Balance & Sigmoid Soft Clipping: Excitation (Glutamate drive) and Inhibition (GABA drive) dynamically balance each other to enforce strict sparsity boundaries. Phase 7c calculates E/I balance where GABA drive dampens Glutamate excitation, applying sigmoid soft clipping to prevent runaway saturation .
Stable Binary Coding & Memory Fast learning of novel patterns overwrites previously learned binary codes (catastrophic interference). Interleaved Sleep Replay & Up/Down-State Selection: Replays novel candidate codes alongside stable “anchor” patterns during offline slow-wave sleep. Phase 6a/6b and Phase 7d execute multi-cycle oscillations (simulated up-states and down-states), interleaving novel candidate patterns with active anchor pools .
Homeostasis & Parameter Stability Binary weights or parameters freeze at boundary extremes (0 or 1), collapsing representation capacity over extended runtime. Saturation Homeostasis & Retrograde Bias Renormalization: Detects boundary saturation streaks and applies sliding-threshold adjustments and mid-target pulls. Phase 6d and Phase 7b track saturation streaks (SATURATION_STREAK_THRESHOLD = 3) at boundary limits (1e-4), applying Anandamide LTD and bias renormalization to gently pull saturated values toward mid-range targets .
Architectural Grouping & Selection Noise or unverified binary patterns are promoted directly into active circuit memory. Guarded Graduation & Critic Snapshot Gating: Candidate representations must survive multiple sleep cycles and pass critic safety checks before permanent promotion. Stage B guarded hypothesis graduation requires candidate codes to survive multi-cycle slow-wave sleep tests, validated against Phase 6b critic snapshot gates .

Key Recommendations for Your DCC Framework & Reference Model

1. Implement Digital Neuromodulation for Dynamic Plasticity

  • Problem in DCCs: Most implementations (e.g., HTM, Sparsey, SPH) rely on static learning rates and fixed \(k\)-winners parameters across all processing steps.
  • BrainStem Food for Thought: Introduce a central state vector of digital neuromodulators . Higher Dopamine and Glutamate can temporarily increase exploration pressure and widen winning context windows . Conversely, higher Serotonin and GABA tighten selection thresholds, increase inhibition, and stabilize established representations .

2. Introduce Offline “Slow-Wave Sleep” Replay Cycles

  • Problem in DCCs: Real-time continuous stream learning frequently suffers from self-generated error propagation or catastrophic interference when novel sparse binary codes overlap with existing ones.
  • BrainStem Food for Thought: Incorporate an offline consolidation phase operating in \(N\)-cycle oscillations (simulating up-states and down-states) . Build an anchor pool of high-stability representations and interleave them with novel candidate codes during sleep replay . Use competitive selection pressure—modulated by GABA/Glutamate ratios—to reinforce valid sparse codes, weaken redundant ones, and prune noise .

3. Enforce E/I Balance & Saturation Homeostasis

  • Problem in DCCs: High-dimensional sparse binary networks can experience parameter saturation or loss of sparsity when processing non-stationary data streams over thousands of cycles.
  • BrainStem Food for Thought: Calculate an explicit Excitation/Inhibition (E/I) balance at every cycle . Monitor boundary saturation streaks when parameters reach extreme values . When saturation occurs alongside flat performance, trigger retrograde dampening (analogous to endocannabinoid Anandamide LTD) to pull saturated weights back toward baseline mid-targets .

4. Establish Guarded Multi-Stage Pattern Graduation

  • Problem in DCCs: Single-pass Hebbian updates can instantly commit noise or transient anomalies to long-term memory.
  • BrainStem Food for Thought: Separate tentative context hypotheses from permanent circuit structures . Require new sparse binary patterns to undergo guarded graduation, verifying that they survive a minimum threshold of slow-wave consolidation cycles and pass a critic snapshot gate before gaining long-term stability .

You can just impose context (for module/neuron/parameter selection) rather than have the system learn it. And then rely on the (linear) map leaning ability of modules,neurons or parameters to adapt the system to the fixed structure imposed by the context system.

I think people dramatically underestimate the ability of local mapping to adapt to imposed structure. And from the get go decide that must make the structure learnable. And then you get all kinds of very complicated schemes that are actually unnecessary.

A viewpoint then is:
You only need to learn the local mappings, you don’t need to learn what selects them.
Of course what selects them must be reasonably sensible, but that is all it need be.

@unikum-sol Thanks for the reply. Although the AI slop portion of your comment is a bit whimsical. I found it very amusing having an LLM committed to your architecture’s context harshly criticize my architecture. I can imagine my LLM and your LLM harshly criticizing each other and explaining that to solve your (possibly non-existent) problems, you only need to adopt my architecture instead :wink:

I’m all for this. I just haven’t got a large enough circuit where global modulation would make sense yet.

Have several different WTA algorithms and none really rely on threshold. It’s just a Top-K sort of cells’ activity, then activating the K cells. The WTA algorithms implicitly model lateral inhibition between cells in a layer. I have a couple other WTA approaches that use regions, topology, and prediction hypotheses.

So far, I haven’t experience catastrophic forgetting. Then again, my systems haven’t been large enough, nor the datasets large enough, where capacity becomes an issue. For the most part, learning new things does not in anyway impact other learned things until the capacity has been reach. Actually, now that i’m thinking about it, i have done a number of RL control problems where exceeded the memory capacity of my circuit.

Our learned representations are graded as you say. However, they are not discrete memories, but imprinted experiences that may eventually coalesce into an interpretable memory. Prior to that the DCC circuits are just learning what is normal, and until the sequence memory crystalizes into common events. New events are easily disposable unless reinforced.

I have realized that some kind of replay mechanism has become necessary, not only for reinforcing memory, but forecasting actions, navigation, and expected events. This has been particularly important for reinforcement learning which relies heavily on large historical playbacks of learning episodes to properly credit historical actions to final outcomes.

Thanks for the thoughtful reply. Fair point on the LLM battle, it really is easy forthese models to get overly dramatic about architectural choices. I totally get your bottom-up engineering approach. It makes complete sense not to add complex global modulation layer when you are still perfecting the individual building blocks and smaller scale dynamics. My point about the Brainstem architecture wasn’t meant to dismiss the elegance of yuor local rules like Top-K inhibition and permanence thresholds. They do a great job handling local stability. But as you scale up and push into RL or complex continuous environments, those local rules naturally hit anceiling once network capacity fills up. That is where subcortical structures really shine. A dedicated Brainstem like layer gives you a systemic off switch to gate plasticity, manage capacity,and trigger ofline replay for credit assignment without polluting the online visual circuits. Seeing you mention that replay mechanisms are becoming necessary for forecasting and credit assignment in your RL setups was awesome because that is exactly where those two worlds meet. Super impressive work on DCC Studio so far.The WASM performance alone is wild. I am definitely looking forward to seeing how your architecture evolves as you scale those circuits up