# SDR questions for image encoding (newbie)

**URL:** <https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674>\
**Category:** Engineering\
**Tags:** encoders, question\
**Created:** [December 14, 2016, 11:44pm UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674 "2016-12-14T23:44:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![naina](https://avatars.discourse-cdn.com/v4/letter/n/919ad9/32.png) [@naina](https://discourse.numenta.org/u/naina)\
**Post date:** [December 14, 2016, 11:44pm UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/1 "2016-12-14T23:44:29Z")

</div>

Hi everyone, I have been a lurker for a while now and reading many of the questions, but now signing up to post my question as I have been working towards my own implementation of Sparse Vectors.I am trying to encode images to enable an SDR to recognize , say cats and dogs - a canonical problem in machine vision.

1. I followed the guidelines in the [encoding data for htm] ([Encoding Data for HTM Systems](https://discourse.numenta.org/t/encoding-data-for-htm-systems/400)) and encoding the pixel levels in each of the `Blue,Green,Red` Channels with `w=50 and n=1000`. Here is the python code for generate the feature vector:  
`cv2.resize(image, (32,32).flatten()`

2. I trained it on cats and dogs image samples and as suggested in the paper,  
`v = int(pixel/256*1000)` -\> value bit  
I set the value bit -\> value + w bit and get recommended sparsity (\<5%),

3. To make predictions I compare a new image against by `or` ing against the SDRs and taking a Jaccard i.e. and\_bits.count(1)/or\_bits.count(1), where and\_bits = query\_sdr & ref\_sdr  
however the prediction results are less than thrilling.

Am I doing this correctly? I seem to have followed the instructions, but no dice ☹  
Can anyone help?

---

<div class="post-metadata">

**Author:** ![jakebruce](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jakebruce/32/996_2.png) [@jakebruce](https://discourse.numenta.org/u/jakebruce)\
**Post date:** [December 15, 2016, 12:18am UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/2 "2016-12-15T00:18:36Z")

</div>

Hi @naina, welcome!

Would it be fair to say your data is not temporal in nature? If so, are you using only the spatial pooler SDR as opposed to the sequence memory SDR? Sequence memory representations are unlikely to be useful for computer-vision-style single-image classification where there is no meaningful time dimension.

If the data is not temporal in nature, HTM may not be quite the right fit for this problem, because HTM is about modelling temporal sequences.

Assuming a sequential version of this problem such as video analysis however, the encoding also strikes me as an issue. I’m assuming you’re doing this according to the following section in the paper:

"8. Encoding Multiple Values

Some applications require multiple values to be  
encoded for a single HTM model. The separate values  
can be encoded on their own and then concatenated to  
form the combined encoding."

If you have 32x32x3 feature vectors where each feature is encoded by 1000 bits, that’s a 3072000-bit input vector? If so, that is probably far too large an input space to learn to classify high level images like dogs and cats, unless you have millions of training samples.

My work involves images, and I’ve found the encoding to be the most important step. The problem with images is that they’re so high-dimensional, the system would need an enormous amount of training data to learn anything useful. So I usually encode images by preprocessing with a standard sparse coding mechanism, like a bank of Gabor filters.

---

<div class="post-metadata">

**Author:** ![naina](https://avatars.discourse-cdn.com/v4/letter/n/919ad9/32.png) [@naina](https://discourse.numenta.org/u/naina)\
**Post date:** [December 15, 2016, 12:31am UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/3 "2016-12-15T00:31:52Z")

</div>

Thanks @jakebruce, I was concerned about the lakc of temporality. But noticed that nuPic Vision had static image encoders. However, it might be possible to add temporality by rotating the image and saving the encodings therein.

> [@jakebruce](#):
>
> Would it be fair to say your data is not temporal in nature? If so, are you using only the spatial pooler SDR as opposed to the sequence memory SDR? Sequence memory representations are unlikely to be useful for computer-vision-style single-image classification where there is no meaningful time dimension.
> 
> If the data is not temporal in nature, HTM may not be quite the right fit for this problem, because HTM is about modelling temporal sequences.
> 
> "8. Encoding Multiple Values
> 
> Some applications require multiple values to beencoded for a single HTM model. The separate valuescan be encoded on their own and then concatenated to form the combined encoding."

Yes indeed I am following the above approach of “extending” bitarrays to form a combined encoding. Thanks for the suggestion of using a Gabor bank, have you tried using any of the convolution approaches that are in vogue now?

Are you able to share any of your approaches (code/papers) so us newbies can learn?  
Thanks again for responding!

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [December 15, 2016, 12:38am UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/4 "2016-12-15T00:38:08Z")

</div>

> [@jakebruce](#):
>
> Would it be fair to say your data is not temporal in nature?

@naina This the important question. From your description, it seems like you’re trying to solve a spatial problem (image classification). HTM is more tuned towards temporal problems.

Here are a few resources that might be informative:

- [HTM suggestions for image recognition](http://lists.numenta.org/pipermail/nupic-hackers_lists.numenta.org/2015-October/004342.html)
- [Vision Object Recognition Using NuPIC](https://github.com/numenta/nupic/wiki/Vision-Object-Recognition-Using-NuPIC)
- [NuPIC for image recognition](http://lists.numenta.org/pipermail/nupic_lists.numenta.org/2015-June/011266.html)
- [Questions about NuPIC Vision](https://discourse.numenta.org/t/questions-about-nupic-vision/487)

---

<div class="post-metadata">

**Author:** ![jakebruce](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/jakebruce/32/996_2.png) [@jakebruce](https://discourse.numenta.org/u/jakebruce)\
**Post date:** [December 15, 2016, 1:36am UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/5 "2016-12-15T01:36:46Z")

</div>

I agree with Matt; temporality is the single relevant question.

But just to answer your question about convolution: yes. The right way to do image classification right now is convolutional neural networks. You can use a pre-trained network like VGG-net and get very nice feature vectors by looking at the intermediate layer representations, and you can binarize these and/or take the top 2% which makes a very good sparse encoding. Gabor filters are just the simplest version of this (the lowest layer of a CNN usually learns Gabor-like filters).

This is not limited to HTM, but these features do work very well as an image encoder for HTM.

---

<div class="post-metadata">

**Author:** ![Sean\_O\_Connor](https://avatars.discourse-cdn.com/v4/letter/s/34f0e0/32.png) [@Sean\_O\_Connor](https://discourse.numenta.org/u/Sean_O_Connor)\
**Post date:** [December 15, 2016, 7:03am UTC](https://discourse.numenta.org/t/sdr-questions-for-image-encoding-newbie/1674/6 "2016-12-15T07:03:25Z")

</div>

I suppose you could use wavelets for feature detection or the SIFT algorithm is very popular at the moment.  
[http://www.inf.fu-berlin.de/lehre/SS09/CV/uebungen/uebung09/SIFT.pdf](http://www.inf.fu-berlin.de/lehre/SS09/CV/uebungen/uebung09/SIFT.pdf)
