# Cortical.io implementation, thoughts and questions

**URL:** <https://discourse.numenta.org/t/cortical-io-implementation-thoughts-and-questions/4704>\
**Category:** Engineering\
**Tags:** nlp, semantic-folding\
**Created:** [October 11, 2018, 11:02pm UTC](https://discourse.numenta.org/t/cortical-io-implementation-thoughts-and-questions/4704 "2018-10-11T23:02:10Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![Paul\_Lamb](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/paul_lamb/32/2725_2.png) [@Paul\_Lamb](https://discourse.numenta.org/u/Paul_Lamb)\
**Post date:** [October 12, 2018, 12:34pm UTC](https://discourse.numenta.org/t/cortical-io-implementation-thoughts-and-questions/4704/3 "2018-10-12T12:34:48Z")

</div>

> [@SporkingIt](#):
>
> The example being used is “apple” appearing in three different contexts: “software”, “fruit” and “desktop”. I can not find any description of how these contexts are deduced or how using one of these contexts to further specify what should be returned in future semantic searches affects the algorithm.
> 
> Am I just blind or is this also part of the proprietary technology? Has anyone any ideas as to how these contexts are deduced and how they are used in semantic searches?

I don’t have any insights into the official algorithm, but I have also developed my own implementation of semantic folding based on the videos that they posted. The way I have approached this task is in addition to generating the word SDRs, also keep track of their frequency. The frequency can be used to produce “weighted SDRs” (non-binary SDR, where the bits have a weight rather than just a zero or one value).

Two weighted SDRs can be compared to generate a weighted overlap score, which allows you to suggest other words which have a lot of overlap and are frequently used. The top results can be used as “contexts” when doing other SDR math.

One other important point is that the words most frequently used in a language tend to add the least amount of uniqueness to the context. This is a result of [Zipfs law](https://discourse.numenta.org/t/mind-blown/3717). This property is very useful to keep in mind when designing “word math” algorithms, since it means for many tasks you can throw out a large percentage of results and focus in on a smaller subset. I bring this up, because you will find that it is relevant to the “contexts” process I described above.

> [@SporkingIt](#):
>
> When it comes to snippet distribution, I claim that it is more important to utilize as much of the space as possible to represent as much semantic information as possible rather than grouping semantically similar snippets close to each other. To this end, I’m using a hashing function to spread snippets in the space and make no attempts to inspect already added snippets.

Wouldn’t the total number of contexts (i.e. bits in the SDR) be the same, or does your hashing function have some scaling property? _(EDIT – NM, I get what you mean – you are essentially using random distribution to map points into a smaller space, and relying on property of SDRs that says a lot of random overlap is virtually impossible)_. Personally I have found that the main advantage to positioning semantically similar snippets close to each other is that you can perform a simple scaling algorithm on the massive original SDRs to produce much smaller working SDRs, so that points close to each other in the large SDR (which are semantically similar to each other) will map to the same point on the smaller SDR.

---

_[View the full topic](https://discourse.numenta.org/t/cortical-io-implementation-thoughts-and-questions/4704)._
