# Does Entropy correctly evaluate whether the SP eﬃciently utilizes all mini-columns?

**URL:** <https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068>\
**Category:** Numenta Theory\
**Tags:** question\
**Created:** [December 15, 2018, 10:19am UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068 "2018-12-15T10:19:59Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![bowen](https://avatars.discourse-cdn.com/v4/letter/b/3be4f8/32.png) [@bowen](https://discourse.numenta.org/u/bowen)\
**Post date:** [December 15, 2018, 10:19am UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/1 "2018-12-15T10:19:59Z")

</div>

We define the entropy that ![image](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/2X/d/da7eff3b67c1e8e4b013cc11455b951d90384c6e.png)

where P(a\_i) is defined by ![image](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/2X/9/988f1583f68aac911d7ada4de3f0b09bdebf00b1.png) , which indicates the average activation frequency of the _i_’th mini-column during _M_ input timesteps.

and the function curve is

![image](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/2X/e/eb1118505a52c769b6418f9e2afc43cddc49bf4d.png)

From the paper "_The HTM Spatial Pooler—A Neocortical Algorithm for Online Sparse Distributed Coding_  
", we know that:

> The SP will have low entropy if a small number of the SP mini-columns are active very frequently and the rest are inactive. Therefore, the entropy metric quantiﬁes whether the SP eﬃciently utilizes all mini-columns.

Then there comes some doubts. It is obviously that when P(a\_i) equals 0.5, the entropy becomes maximum. If we set the activation density to be 2% (i.e. the sparsity should become 2%), while there is some error causing the sparsity to be 50%, then the entropy will be much larger then the correct ones, and we say the SP eﬃciently utilizes all mini-columns. That is not reasonable, isn’t it?

---

<div class="post-metadata">

**Author:** ![marty1885](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/marty1885/32/1891_2.png) [@marty1885](https://discourse.numenta.org/u/marty1885)\
**Post date:** [December 15, 2018, 3:25pm UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/2 "2018-12-15T15:25:44Z")

</div>

Good skeptical thinking!  
I think the assumption is that the SDR density that a SP generates is constant. Under such condition I don’t hink there is a way to exploit the formula.

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [December 17, 2018, 5:03pm UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/3 "2018-12-17T17:03:20Z")

</div>

Without boosting, SP will not efficiently utilize all mini-columns. And yes, we do set a use a constant activation sparsity throughout the process. We don’t change it as time passes or depending on what is being represented.

---

<div class="post-metadata">

**Author:** ![dmac](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/dmac/32/5842_2.png) [@dmac](https://discourse.numenta.org/u/dmac)\
**Post date:** [December 18, 2018, 4:26pm UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/4 "2018-12-18T16:26:47Z")

</div>

It is possible to “normalize” the entropy into the range 0-1. To do this divide by the entropy of the average activation frequence (this is either the hardcoded target freq OR it can be calculated from the data). A result of 1 or 100% indicates maximum utilization, and 0% means the program has serious problems. This normalization makes entropy into a useful debugging tool

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [December 18, 2018, 4:31pm UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/6 "2018-12-18T16:31:25Z")

</div>

> [@dmac](#):
>
> It is possible to “normalize” the entropy into the range 0-1.

Perhaps an interesting technique to fine-tune boosting.

---

<div class="post-metadata">

**Author:** ![ExoBlue](https://avatars.discourse-cdn.com/v4/letter/e/a9adbd/32.png) [@ExoBlue](https://discourse.numenta.org/u/ExoBlue)\
**Post date:** [December 18, 2018, 7:44pm UTC](https://discourse.numenta.org/t/does-entropy-correctly-evaluate-whether-the-sp-e-ciently-utilizes-all-mini-columns/5068/7 "2018-12-18T19:44:30Z")

</div>

Distantly related to this thread, folks have looked at classes of logic circuits that maintain the same number of “0-nodes” and “1-nodes” for stable power consumption.

What might be more relevant – if thinking about entropy as a metric correlated with efficiency – is whether higher layers of cognition follow some sort of Boltzmann distribution in physical count or functional activity.

For example, incoming audio at 16 bits of resolution and a 5 KHz cutoff – i.e. 10K samples per second – shrinks from 20Kbytes per second to about 20 bytes per second if reduced to a single voice speaking.

And, intuitively contrasting today’s speech recognition with that of a couple of decades ago, there’s much more exception-archiving and real-time comparison today.

Which, to this non-expert, is somewhat how a child learns. A couple of dozen or hundred approximate rules to get the gist – and then fine-tuning to the mainstream word and idiom levels of understanding.
