# ANN : Scalar Matrix-Hashing ?encoder?

**URL:** <https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665>\
**Category:** Engineering\
**Created:** [December 12, 2016, 12:29am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665 "2016-12-12T00:29:06Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![mraptor](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/mraptor/32/4469_2.png) [@mraptor](https://discourse.numenta.org/u/mraptor)\
**Post date:** [December 12, 2016, 12:29am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/1 "2016-12-12T00:29:06Z")

</div>

I build Scalar Matrix-Hashing !encoder!.  
Please read the Readme.txt

> **[vsraptor/scalar-hash-encoder](https://github.com/vsraptor/scalar-hash-encoder)**
>
> SDR : Scalar Matrix-Hash encoder. Contribute to vsraptor/scalar-hash-encoder development by creating an account on GitHub.

The interesting thing about it is that it should support as big numbers as possible.  
Additionally you may not need to use Spatial pooler if the hashing works correctly i.e. you can link the Encoder directly with the TM.

The only problem I see is how to guarantee there are no collisions between the different hash functions.  
For this we would need to figure some way to adjust the matrices, based on collisions that happen over time OR figure out way to that from the start !!!

My main reason to post here is to hear from you some ideas on lowering collisions between hashing functions.  
In the current implementation you can check collision statistics :

se.collision\_cnt  
se.avg\_collision

**btw there is no similarity between nearby numbers ! which may disqualify it as Encoder.**

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [December 12, 2016, 11:37pm UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/2 "2016-12-12T23:37:43Z")

</div>

> [@mraptor](#):
>
> btw there is no similarity between nearby numbers ! which may disqualify it as Encoder.

This is true… the first rule of encoding (from [Encoding Data for HTM Systems](https://arxiv.org/abs/1602.05925)) is:

> 1. Semantically similar data should result in SDRs with overlapping active bits.

I’m not sure the results of this encoder will be process-able by HTM systems without that.

---

<div class="post-metadata">

**Author:** ![Sean\_O\_Connor](https://avatars.discourse-cdn.com/v4/letter/s/34f0e0/32.png) [@Sean\_O\_Connor](https://discourse.numenta.org/u/Sean_O_Connor)\
**Post date:** [December 13, 2016, 12:16am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/3 "2016-12-13T00:16:31Z")

</div>

Locality sensitive hashing (LSH) is the way to go:  
[https://en.wikipedia.org/wiki/Locality-sensitive\_hashing](https://en.wikipedia.org/wiki/Locality-sensitive_hashing)  
If you regard LSH as feature (cue) detection and you could extend LSH with some unsupervised learning of those features and you created a read out layer…

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [December 13, 2016, 12:42am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/4 "2016-12-13T00:42:50Z")

</div>

> [@Sean\_O\_Connor](#):
>
> Locality sensitive hashing (LSH) is the way to go

That looks promising. If anyone tries this, please let us all know.

---

<div class="post-metadata">

**Author:** ![mraptor](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/mraptor/32/4469_2.png) [@mraptor](https://discourse.numenta.org/u/mraptor)\
**Post date:** [December 13, 2016, 1:31am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/5 "2016-12-13T01:31:47Z")

</div>

Yes, I figured it out after I posted it ;( … seems I was too quick to post … it seemed too good to be true 😉

---

<div class="post-metadata">

**Author:** ![mraptor](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/mraptor/32/4469_2.png) [@mraptor](https://discourse.numenta.org/u/mraptor)\
**Post date:** [December 13, 2016, 1:39am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/6 "2016-12-13T01:39:01Z")

</div>

Do you know of a hash function that increases collision for similar numbers , as example ?

---

<div class="post-metadata">

**Author:** ![Sean\_O\_Connor](https://avatars.discourse-cdn.com/v4/letter/s/34f0e0/32.png) [@Sean\_O\_Connor](https://discourse.numenta.org/u/Sean_O_Connor)\
**Post date:** [December 13, 2016, 12:57pm UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/7 "2016-12-13T12:57:50Z")

</div>

Ok, you have a list of numbers, in a recomputable way randomly add and subtract those numbers and take the sign of the combination as a binary bit. Hey, did you just destroy all the information in the list? No, actually not. You can say if all the added values were greater than all the subtracted values or not. If you get another bit by using a different random combination that second combination is almost exactly orthogonal to the first. If the list of numbers is sparse in some way the compressive sensing crowd can show you can exactly recreate the list from not so many bits. Obviously if you have two lists that are almost the same they will tend to output the same hash bits. Simple as that really. It isn’t rocket science.  
[http://dsp.rice.edu/sites/dsp.rice.edu/files/cs/CSintro.pdf](http://dsp.rice.edu/sites/dsp.rice.edu/files/cs/CSintro.pdf)

---

<div class="post-metadata">

**Author:** ![Sean\_O\_Connor](https://avatars.discourse-cdn.com/v4/letter/s/34f0e0/32.png) [@Sean\_O\_Connor](https://discourse.numenta.org/u/Sean_O_Connor)\
**Post date:** [December 14, 2016, 1:41am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/8 "2016-12-14T01:41:37Z")

</div>

Maybe this paper is a bit easier:  
[http://www.raeng.org.uk/publications/other/candes-presentation-frontiers-of-engineering](http://www.raeng.org.uk/publications/other/candes-presentation-frontiers-of-engineering)

A simpler viewpoint is that if the input data is sparse the full information content of that data can be captured with a moderate number of random correlations. It is never going to be as compressive as say Jpeg but it makes less assumptions about the data and is easy to do.

For LSH if the data vector has 65536 (2^16) elements and you want 65536 hash bits you have a slight problem there. You would need to do n_n (65536_65536) (±) operations, kind of slow. You can reduce that down to n(ln n) 65536\*16 (±) operations by doing (recomputable) random sign flipping of the input data followed by a Walsh Hadamard transform.

There is an old joke that everything in AI is dot product calculations, and so it is with the above when you look into it.

---

<div class="post-metadata">

**Author:** ![brev](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/brev/32/5548_2.png) [@brev](https://discourse.numenta.org/u/brev)\
**Post date:** [April 19, 2019, 2:38am UTC](https://discourse.numenta.org/t/ann-scalar-matrix-hashing-encoder/1665/9 "2019-04-19T02:38:27Z")

</div>

FYI, I think I ended up getting something working along these lines.

> [@NEW: SimHash Distributed Scalar Encoder (SHaDSE) - DEPRECATED](https://discourse.numenta.org/t/new-simhash-distributed-scalar-encoder-shadse/5860):
>
> Introducing: SimHash Distributed Scalar Encoder (SHaDSE) A [Locality-Sensitive Hashing](https://en.wikipedia.org/wiki/Locality-sensitive_hashing) approach towards encoding semantic data into [Sparse Distributed Representations](https://numenta.com/neuroscience-research/sparse-distributed-representations/), ready to be fed into an [Hierarchical Temporal Memory](https://numenta.com/machine-intelligence-technology/), like [NuPIC](https://github.com/numenta/nupic) by [Numenta](https://numenta.com). This uses the [SimHash](https://en.wikipedia.org/wiki/SimHash) algorithm to accomplish this. LSH and SimHash come from the world of nearest-neighbor document similarity searching. This encoder is sibling with the original [Scalar Encoder](https://github.com/numenta/nupic/blob/master/src/nupic/encoders/scalar.py), and the [Random Distributed Scalar Encoder](https://github.com/numenta/nupic/blob/master/src/nupic/encoders/random_distributed_scalar.py) (RDSE). T…

thanks.
