# Exciting potentials with HTM agents in OpenAI Gym

**URL:** <https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710>\
**Category:** Engineering\
**Created:** [October 16, 2019, 1:59pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710 "2019-10-16T13:59:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![marty1885](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/marty1885/32/1891_2.png) [@marty1885](https://discourse.numenta.org/u/marty1885)\
**Post date:** [October 16, 2019, 1:59pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/1 "2019-10-16T13:59:04Z")

</div>

I’m still working on this. But I’m too excited to hold back myself and want to share the results right now.

Its my graduation project and I’m building RL agents using HTM algorithms. To my surprise, HTM works rather well (comparing to DQN and A2C) in environments with dense enough rewards. HTM can learn how to act in just a few episodes. But the learning seems to collapse after, say 200 training loops. After that HTM just doen’t know how to act.

(Figure: A HTM agent in the CartPole-v1 environment and getting high rewards very early.)

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/2X/b/b0870da47d3df2571164ce8d637f474e7e025c15.jpeg)

I’ll release the source code once it’s ready.

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [October 16, 2019, 2:35pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/2 "2019-10-16T14:35:47Z")

</div>

[![](https://media.giphy.com/media/xxAsO5MtariRW/giphy.gif) ](https://media.giphy.com/media/xxAsO5MtariRW/giphy.gif)

---

<div class="post-metadata">

**Author:** ![marty1885](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/marty1885/32/1891_2.png) [@marty1885](https://discourse.numenta.org/u/marty1885)\
**Post date:** [October 17, 2019, 3:06pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/3 "2019-10-17T15:06:21Z")

</div>

Now I can get better and more stable rewards. But the network still eventually collapses. ☹

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/numenta/original/2X/f/f73bb0eebc35b150198c50d18d52a89a6f479f04.png)

---

<div class="post-metadata">

**Author:** ![rhyolight](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/rhyolight/32/3922_2.png) [@rhyolight](https://discourse.numenta.org/u/rhyolight)\
**Post date:** [October 17, 2019, 4:44pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/4 "2019-10-17T16:44:16Z")

</div>

Has anyone else here used the OpenAI Gym platform?

---

<div class="post-metadata">

**Author:** ![dmac](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/dmac/32/5842_2.png) [@dmac](https://discourse.numenta.org/u/dmac)\
**Post date:** [October 19, 2019, 11:35pm UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/5 "2019-10-19T23:35:04Z")

</div>

I have. It has a variety of reinforcement learning games of varying difficulties. All of the games have the same API, so applying an AI to many different games is easy. It’s nice to work with and experiment with, though I haven’t solved any of them yet.

---

<div class="post-metadata">

**Author:** ![marty1885](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.numenta.org/marty1885/32/1891_2.png) [@marty1885](https://discourse.numenta.org/u/marty1885)\
**Post date:** [October 20, 2019, 4:12am UTC](https://discourse.numenta.org/t/exciting-potentials-with-htm-agents-in-openai-gym/6710/6 "2019-10-20T04:12:29Z")

</div>

Never mind. I found why the agent goes nuts after a while. It turns out that the environment I use (CartPole-v1) requires the agent to alternate between sending commands going right and left to go slow. But HTM is bad at that.
