P-11 // Started

Rotary Double Inverted Pendulum

A rotary double inverted pendulum with slip rings and CAN-connected sensors, built to swing up through all four equilibria with reinforcement learning and to explore MPC as an alternative.

CAD complete ยท Hardware build in progress
  • Reinforcement learning
  • MPC
  • Nonlinear control
  • STM32H7
  • Raspberry Pi 5
  • CAD
Open repository
Section render of the complete rotary double inverted pendulum assembly, with the rotating arm and both pendulum links

This project is currently in progress. The mechanical design is finished and I am building the hardware. The controller is still to be developed.

Goal

The goal is a controller that swings the double pendulum up and moves it through all four equilibrium positions as a transition control: both links down, both up and the two mixed configurations in which one link points up and the other down. The main approach is reinforcement learning, with a policy trained in simulation and then deployed on the real system. I also want to explore model predictive control (MPC) on the same hardware, so the two methods can be compared on the same task.

The system is strongly nonlinear and unstable at the upper equilibria, which makes it a demanding benchmark for both learned and optimisation-based control.

Approach: I plan to implement everything myself, from the mechanical design to the simulation, the reinforcement-learning pipeline and the MPC, without copying designs or setups from other projects.

Hardware

The structure looks large in the renders, but it is compact: 45 cm high, with two 20 cm pendulum links and a base radius of 10 cm. Three encoders measure the arm and both pendulum angles. The motor and the sensors communicate over CAN. An STM32H7 is intended for the real-time part and a Raspberry Pi 5 for the higher-level computation. Slip rings carry power and signals through the rotating joints so cables cannot tangle during unlimited rotation.

Planned workflow

Roadmap
StagePlanned work
1. DesignComplete the CAD model of the base, rotating arm and pendulum links.
2. BuildAssemble the mechanics, slip rings, encoders and CAN electronics.
3. ModelDerive and identify the dynamics and build a simulation.
4. LearnTrain a reinforcement-learning policy for swing-up and transitions between all four equilibria.
5. CompareExplore MPC and compare it with the learned policy on the real hardware.

Results, simulation plots and hardware footage will be added as the project progresses. The source is on GitHub.