# Ogma's Sparse Predictive Hierarchies

**URL:** <https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176>\
**Category:** Tangential Theories\
**Created:** [October 7, 2022, 3:27pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176 "2022-10-07T15:27:37Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![cezar\_t](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/cezar_t/32/4451_2.png) [@cezar\_t](https://discourse.numenta.org/u/cezar_t)\
**Post date:** [October 7, 2022, 3:27pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/1 "2022-10-07T15:27:37Z")

</div>

The purpose of this topic is not only about drawing attention towards Ogma’s ideas/framework which currently is named SPH (from Sparse Predictive Hierarchies) but mostly to discuss how to apply SPH to SDR-based processing, and I’m thinking at both:

- HTM tools, specifically Spatial Pooler and Temporal Memory as encoder/decoder
- Associative Memory tools as in Diadic/Triadic memory and possibly any other bitpair maps.

Why?  
One reason is Ogma already produced some [interesting results](https://www.youtube.com/c/OgmaAi) with quite limited hardware like Raspberry Pi-s and even microcontrollers,  
The other reason being HTM feels kind of stuck at TM and SP - they seem to work but I haven’t seen very explicit proposals to further expand (or assemble) these basic blocks in more complex architectures.

To start with here are a couple links detailing the core concepts in SPH:

- [SPH Presentation](https://raw.githubusercontent.com/ogmacorp/OgmaNeo2/master/SPH_Presentation.pdf)  
(30 slides as PDF)
- [A paper draft](https://raw.githubusercontent.com/ogmacorp/OgmaNeo2/master/OgmaNeo2_Whitepaper_DRAFT.pdf)
- And the [github repository](https://github.com/ogmacorp/OgmaNeo2) with code and references to the above documents.

* * *

One main difference between Ogma’s system vs HTM is their is not based on SDRs as data representation, but on a similar, yet very different structure called CSDRs. I won’t delve into differences, because the papers above are much more clear, and what I am going to propose here is using SDRs instead of Ogma’s CSDRs in a SPH-like system.

* * *

One of the most interesting chart describing SPH is at the page 5 on the paper, or page 14 of the slide presentation.  
I don’t know how to pull out an image from a PDF, so you’ll have to look there to make sense of what follows:

That image contains a layered stack of encoder-decoder pairs, layer 1 being the bottom-most encoder-decoder pair and layer N (usually N=3 or greater) is placed at the top of the SPH hierarchy

From what @ericlaukien kindly explained, there is no inherent constraint on what “encoder” and “decoder” blocks are made of, they experimented with a vast array/kinds of encoders, while for decoders they used mainly a (relatively) simple logistic regression.

One key feature (not obvious in the ladder schematic) is  
each upper layer operates at half the time rate of its underlying layer, (mostly 1/2 time steps). That means increasing number of layers expands the time span of the whole system, without a linear increasing in computing costs. They call this “exponential memory”, in the sense each upper layer “sees” changes over twice the time span of the layer below it.

* * *

Now how can SPH architecture could be build with HTM bricks:

- Use a Spatial Pooler as encoder.
- Use a variation of Temporal Memory as decoder.

Will further discuss how TM needs to be altered in order to be usable in a SPH, because by default TM predicts its own next input while in SPH it has to predict the future underlying encoder output(s)

Without further ado one should also notice that a triadic memory might be good decoder too, while as encoder can be tested a FlyHash encoder that simply “translates” X input bits into X/2 output bits.

---

<div class="post-metadata">

**Author:** ![JarvisGoBrr](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jarvisgobrr/32/6147_2.png) [@JarvisGoBrr](https://discourse.numenta.org/u/JarvisGoBrr)\
**Post date:** [October 7, 2022, 5:41pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/2 "2022-10-07T17:41:24Z")

</div>

It looks like fun to play around with robots running these kinds of algorithms.  
we should try to build stuff like this, even if just robots inside simulated game-like environments.

---

<div class="post-metadata">

**Author:** ![JarvisGoBrr](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jarvisgobrr/32/6147_2.png) [@JarvisGoBrr](https://discourse.numenta.org/u/JarvisGoBrr)\
**Post date:** [October 7, 2022, 6:31pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/3 "2022-10-07T18:31:23Z")

</div>

I once made a very rudimentary toy self-driving car which was kinda similar but used a stack of spatial poolers.

[![](https://img.youtube.com/vi/86z1zQEoZk0/hqdefault.jpg "bot learning to drive.") ](https://www.youtube.com/watch?v=86z1zQEoZk0)

[![](https://img.youtube.com/vi/Z2vL9--WSFU/maxresdefault.jpg "Reinforcement learning test") ](https://www.youtube.com/watch?v=Z2vL9--WSFU)

---

<div class="post-metadata">

**Author:** ![cezar\_t](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/cezar_t/32/4451_2.png) [@cezar\_t](https://discourse.numenta.org/u/cezar_t)\
**Post date:** [October 8, 2022, 9:44pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/4 "2022-10-08T21:44:31Z")

</div>

That’s cool,  
It would be interesting to attack gym’s car racing environment which presents the challenge of being image based.  
Well you can handcraft some “forward looking lasers” to get similar results as in these videos, but it would be more interesting to work more directly with the screen image itself.

---

<div class="post-metadata">

**Author:** ![JarvisGoBrr](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jarvisgobrr/32/6147_2.png) [@JarvisGoBrr](https://discourse.numenta.org/u/JarvisGoBrr)\
**Post date:** [October 8, 2022, 11:34pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/5 "2022-10-08T23:34:44Z")

</div>

I’d have to downscale the image and apply a highpass filter to get it to run at realtime speeds but it should work.

---

<div class="post-metadata">

**Author:** ![jacobeverist](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jacobeverist/32/6736_2.png) [@jacobeverist](https://discourse.numenta.org/u/jacobeverist)\
**Post date:** [October 9, 2022, 12:53am UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/6 "2022-10-09T00:53:53Z")

</div>

@cezar_t

SPH, HTM, and Sparsey (official [website](http://sparsey.com/)/[code](https://github.com/Neurithmic-Systems/V-to-mu_Demo)) are the pioneering exemplars for what I call Binary Pattern Neural Networks (BPNNs), with Sparsey being the first. I describe BPNNs below in my draft paper (which I have yet to complete).

> BPNNs generally have the following properties: 1) they receive input in the form of binary vectors, 2) they use a form of Winner-Take-All (WTA) computation for selecting the neurons to activate, and 3) the neurons have a binary activation function for output. The implementations of BPNNs differ in how neuron activation is implemented, how the network learns, and how the network is architected. BPNNs are not to be confused with Binary Neural Networks (BNNs) [?], which are traditional ANNs but with activation functions that that transform the underlying scalar weights and neuron states into binary. Unlike BNNs, BPNNs natively operate on binary states.

Given the previous discussion on [clusterons](https://discourse.numenta.org/t/gradient-clusteron-visualization/10125/7), I think including dendrites as first-order computational objects should also fall into this umbrella definition. The power of research in this area is in the clear visual exploration of distributed representation and computation of discrete information packets. This is in contrast to the ANN approach which linearizes computation and embeds information into vector spaces. I’m starting to strongly believe the latter’s availability of strong existing linear algebra computational and math tools is severely inhibiting scientific advancements in artificial cognition.

What’s missing for BPNNs, and why HTMs seem to have been stalled, is a clear theoretical framework for how information is represented and transformed through BPNN networks. Without that theory, trying to connect spatial poolers and temporal sequence memories together to create some effect is just like wiring blackboxes together to see what happens. When you fail, there’s no way to understand why you fail or how to improve it without that theoretical understanding.

I think I have part of this theory, but it’s still a long way from explaining what’s happening, and how to get desired effects.

---

<div class="post-metadata">

**Author:** ![thanh-binh.to](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/thanh-binh.to/32/1793_2.png) [@thanh-binh.to](https://discourse.numenta.org/u/thanh-binh.to)\
**Post date:** [October 9, 2022, 8:18am UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/7 "2022-10-09T08:18:40Z")

</div>

@JarvisGoBrr your self-driving car looks nice!  
Do you use HTM + RL?

Personally I found SPH is very powerful frameworks which is continuously developed by Ogma: high performance and very fast.

One thing I find much interesting is imager encoder/decoder of AOgmaNeo, which allows us to check the potentials of HTM for image prediction and classification.

In my experiments HTM works perfectly with CSDR for both image prediction and classification!

---

<div class="post-metadata">

**Author:** ![JarvisGoBrr](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jarvisgobrr/32/6147_2.png) [@JarvisGoBrr](https://discourse.numenta.org/u/JarvisGoBrr)\
**Post date:** [October 9, 2022, 1:01pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/8 "2022-10-09T13:01:03Z")

</div>

I only use spatial poolers for the self driving car.

one is an input encoder, which is decoded into preddicted reward and action by other two spoolers.

---

<div class="post-metadata">

**Author:** ![thanh-binh.to](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/thanh-binh.to/32/1793_2.png) [@thanh-binh.to](https://discourse.numenta.org/u/thanh-binh.to)\
**Post date:** [October 9, 2022, 5:10pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/9 "2022-10-09T17:10:10Z")

</div>

@JarvisGoBrr  
How can you calculate rewards?  
Which RL do you use?  
What are 2 actions? Velocity and car orientation in the driving direction?

---

<div class="post-metadata">

**Author:** ![JarvisGoBrr](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jarvisgobrr/32/6147_2.png) [@JarvisGoBrr](https://discourse.numenta.org/u/JarvisGoBrr)\
**Post date:** [October 9, 2022, 5:25pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/10 "2022-10-09T17:25:15Z")

</div>

I dont know which algorithm I’m using, I just made it up.

First primary input pooler that learns to represent the visual input.

supervisory reward is 0 if car is in road and -1 if its touching the border.

I train a spatial pooler to decode the visual input into a reward prediction but its being trained on its own prediction plus the real reward times an adaptation constant.

This leads it to a “reward smearing” over space and time, that turns a sparse reward into a smoothly varying dense reward that is more negative close to edges of the road and fades off as the car gets further away.

I use this dense reward to modulate the learning of a second pooler that decodes the input into a left-right action.

if dense reward increased relative to previous time step, reinforce the previous action taken otherwize forget the action.

---

<div class="post-metadata">

**Author:** ![thanh-binh.to](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/thanh-binh.to/32/1793_2.png) [@thanh-binh.to](https://discourse.numenta.org/u/thanh-binh.to)\
**Post date:** [October 9, 2022, 6:38pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/11 "2022-10-09T18:38:42Z")

</div>

@JarvisGoBrr ok and thank  
AOgmaNeo learns very quickly and the car runs some rounds without any problem!

---

<div class="post-metadata">

**Author:** ![cezar\_t](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/cezar_t/32/4451_2.png) [@cezar\_t](https://discourse.numenta.org/u/cezar_t)\
**Post date:** [October 10, 2022, 6:52pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/12 "2022-10-10T18:52:45Z")

</div>

@jacobeverist your paper sounds interesting. Do you have any draft public?

When I googled “BPNNs” top results are about Back Propagated Neural Networks which usually means ordinary deep networks.

That’s a potential of confusion.

---

<div class="post-metadata">

**Author:** ![cezar\_t](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/cezar_t/32/4451_2.png) [@cezar\_t](https://discourse.numenta.org/u/cezar_t)\
**Post date:** [October 10, 2022, 8:50pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/13 "2022-10-10T20:50:26Z")

</div>

@thanh-binh.to did you used RL or “teacher learning” as in the featured youtube videos?

---

<div class="post-metadata">

**Author:** ![jacobeverist](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jacobeverist/32/6736_2.png) [@jacobeverist](https://discourse.numenta.org/u/jacobeverist)\
**Post date:** [October 11, 2022, 4:08pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/14 "2022-10-11T16:08:31Z")

</div>

@cezar_t I don’t think that acronym is very commonly used given that nearly all deep neural networks use back propagation for learning these days. Do you have a better name?

Sadly, my draft is under corporate lock&key at the moment so i’m unable to release the partial work without going through a release process. I want it to be complete before I do that so I don’t have to go through it again.

The paper actually focuses mostly on encoders and how to build them which is a sadly neglected but essential topic.

---

<div class="post-metadata">

**Author:** ![thanh-binh.to](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/thanh-binh.to/32/1793_2.png) [@thanh-binh.to](https://discourse.numenta.org/u/thanh-binh.to)\
**Post date:** [October 11, 2022, 6:44pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/15 "2022-10-11T18:44:43Z")

</div>

> [@cezar\_t](#):
>
> @thanh-binh.to did you used RL or “teacher learning” as in the featured youtube videos?

@cezar_t i do not know about this video.  
In Ogmaneo they use actor-critics algorithm

---

<div class="post-metadata">

**Author:** ![cezar\_t](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/cezar_t/32/4451_2.png) [@cezar\_t](https://discourse.numenta.org/u/cezar_t)\
**Post date:** [October 12, 2022, 9:28am UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/16 "2022-10-12T09:28:27Z")

</div>

@jacobeverist Thanks. Then I assume the paper isn’t just a review of BPNNs,. About acronym collision I just noticed it , could be “representations” instead of “patterns”, or ML instead of NN.  
You have a better perspective on how important this detail is.  
PS. BRNN is taken by the less notorious Bidirectional Recurrent NNs which I haven’t heard of yet, but sounds interesting.

* * *

@thanh-binh.to [interesting results](https://www.youtube.com/c/OgmaAi) link in first message here is their youtube channel demonstrations.  
The most impressive ones feature either some form of imitation learning, or some simple kind of path memorization followed by goal setting.  
In both the robot is first manually driven through an environment, which is not what mainstream RL is about.  
However their system can also be set to work in RL settings. I played only with the cartpole example which if time-scaled to real time it isn’t performing as well as the much more complex raspbery pi RC car in the park alley with imitation learning.

That’s why I asked what kind of results are you talking about.

---

<div class="post-metadata">

**Author:** ![thanh-binh.to](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/thanh-binh.to/32/1793_2.png) [@thanh-binh.to](https://discourse.numenta.org/u/thanh-binh.to)\
**Post date:** [October 15, 2022, 6:25pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/17 "2022-10-15T18:25:43Z")

</div>

> [@cezar\_t](#):
>
> hat’s why I asked what kind of results are you talking

I am speaking about car racing demo!

---

<div class="post-metadata">

**Author:** ![dmac](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/dmac/32/5842_2.png) [@dmac](https://discourse.numenta.org/u/dmac)\
**Post date:** [October 17, 2022, 3:35pm UTC](https://discourse.numenta.org/t/ogmas-sparse-predictive-hierarchies/10176/18 "2022-10-17T15:35:34Z")

</div>

Another peice of prior art is the “cerebellar model articulation controller” which almost meets your definition (IIRC it does not use a competition but it does incorporate sparsity). Its from the '70s. The authors analysis of how sparsity affects it is basically correct but much less math-formalized than the state-of-the-art theories.

> **[Cerebellar model articulation controller](https://en.m.wikipedia.org/wiki/Cerebellar_model_articulation_controller)**
>
> The cerebellar model arithmetic computer (CMAC) is a type of neural network based on a model of the mammalian cerebellum. It is also known as the cerebellar model articulation controller. It is a type of associative memory.
> The CMAC was first proposed as a function modeler for robotic controllers by James Albus in 1975 (hence the name), but has been extensively used in reinforcement learning and also as for automated classification in the machine learning community. The CMAC is an extension of...
