# Open-source text SDR encoder?

**URL:** <https://discourse.numenta.org/t/open-source-text-sdr-encoder/12127>\
**Category:** Engineering\
**Tags:** question\
**Created:** [November 28, 2025, 7:38pm UTC](https://discourse.numenta.org/t/open-source-text-sdr-encoder/12127 "2025-11-28T19:38:55Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![sean](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/sean/32/6652_2.png) [@sean](https://discourse.numenta.org/u/sean)\
**Post date:** [December 11, 2025, 4:29am UTC](https://discourse.numenta.org/t/open-source-text-sdr-encoder/12127/3 "2025-12-11T04:29:10Z")

</div>

Thx for the pointers.

I am looking for guidance on how to optimize the text SDR enocoding for optimal text retrieval properties. My intuition is: maximize entropy of the SDR cells aka every cell in the SDR should represent very unique / niche semantic concepts and have at best no overlapping words they represent (even if they are close neighbors).

Just found these related threads:

- [Cortical.io implementation, thoughts and questions](https://discourse.numenta.org/t/cortical-io-implementation-thoughts-and-questions/4704)
- [Cortical.io encoder algorithm docs](https://discourse.numenta.org/t/cortical-io-encoder-algorithm-docs/707)
- [Words to SDR?](https://discourse.numenta.org/t/words-to-sdr/3660)
- [How can I create an Artificial Intelligence system that can learn new words that are out of their initial vocabulary?](https://discourse.numenta.org/t/how-can-i-create-an-artificial-intelligence-system-that-can-learn-new-words-that-are-out-of-their-initial-vocabulary/3758)

---

_[View the full topic](https://discourse.numenta.org/t/open-source-text-sdr-encoder/12127)._
