# Trouble using Category encoder C as special flag and with string datatype

**URL:** <https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668>\
**Category:** NuPIC\
**Created:** [May 31, 2016, 10:51am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668 "2016-05-31T10:51:19Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [May 31, 2016, 10:51am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/1 "2016-05-31T10:51:19Z")

</div>

**Exception: E10002: Exiting due to receiving too many models failing from exceptions (6 out of 6).**  
**assert types[self.\_resetIdx] == FieldMetaType.integer**

Q1) I am getting this assertion error whenever I am using C (category encoder) as a special flag in data set for swarming ? And when I change the special flag to R or S then it is working fine .  
Q2) I am getting the same assertion error while using string as datatype for any input field .

Can anybody help me in resolving these issues ?

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [May 31, 2016, 2:27pm UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/2 "2016-05-31T14:27:33Z")

</div>

Please show us:

- swarm description JSON file
- small sample of data CSV, including headers

---

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [June 1, 2016, 5:15am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/3 "2016-06-01T05:15:23Z")

</div>

I am attaching below sample of my dataset ![](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/1X/3a3ab8ca4219df920d5b784479b9aa2053704574.PNG)

And regarding swarmdescription file . Actually I have automated everything so It is automatically reading everything from csv file and copy them to the required dictionary which I have checked correct only. As a result making it easy to run swam over any type of dataset

---

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [June 1, 2016, 6:44am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/4 "2016-06-01T06:44:08Z")

</div>

Note that In 2nd column 1 represents N(north) and 0 represents S(south). So if I change this 1 to N and 0 to S with string as its datatype and no special flag then swarm ran successfully but it is not giving the expected results . Results are highly deviating from the actual values. I am not able to understand the reason .

Below I am attaching the output file and altMap ( prediction1 represent prediction after 1 timestamp and prediction2 represents prediction after 2 timestamp)

![](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/1X/3bdbe8b1ddd495eb34a5b00961da53b929195b6b.PNG)  
 ![](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/1X/432e2f6c1fa492f1aaf96adc61b4c4bc9fb752c3.PNG)  
 ![](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/1X/85b4b997c4b429dd799be7ba9282876d4cbc6df7.PNG)

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [June 1, 2016, 1:25pm UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/5 "2016-06-01T13:25:30Z")

</div>

It looks like you have marked two data fields as `R` for `reset`, but this doesn’t make sense. The values in those fields don’t look like they are actually resetting sequences properly, either. See [NuPIC Input Data File Format](https://github.com/numenta/nupic/wiki/NuPIC-Input-Data-File-Format#row-3-flags) for details about how reset works.

> [@priyankyadav31](#):
>
> And regarding swarmdescription file . Actually I have automated everything so It is automatically reading everything from csv file and copy them to the required dictionary which I have checked correct only. As a result making it easy to run swam over any type of dataset

Ok, but there still must be a way to get that dict and print it to the screen during runtime. It probably isn’t right. I can’t even tell which field you are predicting without looking at the swarm file.

> [@priyankyadav31](#):
>
> Note that In 2nd column 1 represents N(north) and 0 represents S(south). So if I change this 1 to N and 0 to S with string as its datatype and no special flag then swarm ran successfully but it is not giving the expected results . Results are highly deviating from the actual values. I am not able to understand the reason .

The swarming process creates model parameters, then those model parameters are used to create a model. What is your process? Are you swarming every time you run? I think you have something wrong at the core of your process… if I could see your code, it would help me debug it.

---

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [June 2, 2016, 4:32am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/6 "2016-06-02T04:32:01Z")

</div>

> [@rhyolight](#):
>
> It looks like you have marked two data fields as R for reset, but this doesn’t make sense. The values in those fields don’t look like they are actually resetting sequences properly, either. See NuPIC Input Data File Format for details about how reset works.

Actually I am not getting the exact meaning of using these special meaning and how to use them even after going through the source that you have mentioned . So it would be very helpful to me if you can elaborate when and for which type of field what special flag should be used?

> [@rhyolight](#):
>
> Ok, but there still must be a way to get that dict and print it to the screen during runtime. It probably isn’t right. I can’t even tell which field you are predicting without looking at the swarm file.

Here column A,B,C,D,E are the inputs and column F is the output field to be predicted.  
And column G is the prediction of column F for 1 timestamp ahead and column H is the prediction of column F for 2 timestamp ahead?

![](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/1X/6f93f533e612b3ca676889935b3266ad294888bc.PNG)

---

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [June 2, 2016, 5:46am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/7 "2016-06-02T05:46:09Z")

</div>

[https://drive.google.com/folderview?id=0BxYSZ7cosxJxR3lwWjhVY3dVLXM&usp=sharing](https://drive.google.com/folderview?id=0BxYSZ7cosxJxR3lwWjhVY3dVLXM&usp=sharing)

I am attaching this link containing the required files

---

<div class="post-metadata">

**Author:** ![priyankyadav31](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/priyankyadav31/32/290_2.png) [@priyankyadav31](https://discourse.numenta.org/u/priyankyadav31)\
**Post date:** [June 14, 2016, 7:29am UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/8 "2016-06-14T07:29:17Z")

</div>

I have not received any replies since 12 days . Please let me know if it is in an inappropropriate format ?

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [June 15, 2016, 2:51pm UTC](https://discourse.numenta.org/t/trouble-using-category-encoder-c-as-special-flag-and-with-string-datatype/668/9 "2016-06-15T14:51:53Z")

</div>

> [@priyankyadav31](#):
>
> I have not received any replies since 12 days . Please let me know if it is in an inappropropriate format ?

Sorry I did not get back to you sooner, but I’ve been traveling and been on vacation for the past week. Before I look at your code, I’d like to continue our conversation.

> [@priyankyadav31](#):
>
> Actually I am not getting the exact meaning of using these special meaning and how to use them even after going through the source that you have mentioned . So it would be very helpful to me if you can elaborate when and for which type of field what special flag should be used?

The `R` (reset) flag is outside of your data. It is a way for you to indicate that a sequence has ended. Every time NuPIC sees a `1` in that field, it indicates that some sequence in the data has completed and another will start. This should correspond with some logical break in input data sequences. For example, when analyzing vehicle movements, sequences might be reset after the engine turns off.

The `S` (sequence) flag has the same purpose, but a different method. Instead of indicating when a sequence should reset, this is a convenience to allow you to add an id to the input data to identify different sequences (track ids in GPS tracks, for example). Whenever this value changes in the data, NuPIC assumes a new sequences is occurring.

> [@priyankyadav31](#):
>
> Here column A,B,C,D,E are the inputs and column F is the output field to be predicted.And column G is the prediction of column F for 1 timestamp ahead and column H is the prediction of column F for 2 timestamp ahead?

This is a row of input, not what we need to look at. We need to see the model parameters. These are used when creating a model. It is the dict you pass into a model when it is created. That is what we need to print out and see:

```python
from pprint import pprint
pprint(model_params)

```

Use the [`pprint`](https://docs.python.org/2/library/pprint.html) module to get a nicely formatted output.
