Abstract
Soft continuum robots have been accepted as a promising category of biomedical robots, accredited to the robots’ inherent compliance that makes them safely interact with their surroundings. In its application of minimally invasive surgery, such a continuum concept shares the same view of robotization for conventional endoscopy/laparoscopy. Different from rigid-link robots with accurate analytical kinematics/dynamics, soft robots encounter modeling uncertainties due to intrinsic and extrinsic factors, which would deteriorate the model-based control performances. However, the trade-off between flexibility and controllability of soft manipulators may not be readily optimized but would be demanded for specific kinds of modeling approaches. To this end, data-driven modeling strategies making use of machine learning algorithms would be an encouraging way out for the control of soft continuum robots. In this article, we attempt to overview the current state of kinematic/dynamic model-free control schemes for continuum manipulators, particularly by learning-based means, and discuss their similarities and differences. Perspectives and trends in the development of new control methods are also investigated through the review of existing limitations and challenges.
1 Introduction
Bioinspired by snakes, elephant trunks, and octopus tentacles, continuum robots are designed to structurally mimic their inherent dexterity and adaptability (Webster and Jones, 2010). In contrast to conventional rigid-link manipulators, “continuum” mechanisms leverage a series of continuous arcs without a skeletal structure to produce a bending motion (). Such design initially focuses on large-scale grasping, locomotion, and positioning in industrial applications () or even urban search and rescue operations in confined environments (). The trade-off between high flexibility and low payload induces strict structural requirements. Gradually, with the reduced scale of continuum robots, the concerns are also diverted to the delicate steering of the slim robot body. The flexible characteristics of continuum robots are appropriate for surgical field applications. Enabling infinite degree-of-freedom (DoF) manipulations within small scales, continuum robots endow the target with flexible access and the patient with less invasion (). Moreover, pliable interventional devices with broad-range functions, such as catheters (; Wang et al., 2018), afford much inspiration to the development of robotic continuum manipulators. Besides the mechanism of a robot, proper controllers and corresponding sensors are also necessary to guarantee accurate control performance.
Conventional rigid-link robots could be controlled directly by the commands of motors on each joint. Once the joint angles and the link lengths are available, the pose of all points, including the end-effector, can be fully determined. Different from rigid-link mechanisms which have definite kinematic mapping, soft robots meet great challenges in accurate analytical modeling, considering the nonlinear deformation induced by the actuation, material elasticity, and susceptibility to contacts with surroundings. Generally speaking, the actuation mechanism of continuum robots (; Yip and Camarillo, 2014; ; ; ; ; ) can be categorized as intrinsic and extrinsic (), based on the location of the actuator. The intrinsic mechanism means that the actuators are located inside and form as part of the mechanism (). One example could be pneumatic-driven robots (Figures 1A–E), whose deformation is induced by the inflation of internal elastic chambers. Extrinsic mechanisms use external components to distort the robot body (Figures 1F,G), such as tendons/cables dragged by motors. Due to the high compliance of continuum manipulators, constraints imposed by obstacle interactions may deform the robot body into undesired shapes regardless of the actuation status. These effects are generally difficult to completely sense and feed back into the model, leading to unstable behaviors. In addition, the individual variation also leads to modeling uncertainties. For example, even with the same design prototype, different degrees of fabrication errors may require repetitive parameter tuning for all robots and their actuators (). The characteristics of a specific robot may also change (e.g., wear-out effects) over the course of time.
FIGURE 1
Machine learning approaches provide a promising way out for the control of continuum robots. As the controller or inverse kinematic mapping is identified by experimental sensory data, people also title it as data-driven control. Sometimes, to distinguish it from the quantitative modeling which parameterizes the system using compact representations (e.g., differential equations) (
Various machine learning techniques have the potential or have been validated for continuum robot control in existing works. In this article, we aim to summarize and discuss their representative implementations. Different from previous reviews such as the one by
2 Iteration-Based Kinematic Model-Free Control
In model-based controllers, the kinematic model can be decomposed into two mappings (
FIGURE 2

Mapping relationship between the three spaces (i.e., : actuation space, : configuration space, and : task space) in soft robot control.
2.1 Optimization-Based Jacobian Matrix Estimation
Yip and Camarillo (2014), Yip and Camarillo (2016), and Yip et al. (2017) proposed a series of model-free closed-loop controllers based on the optimal control strategy. Different from deducing the Jacobian matrix regarding the analytical model (
Assume that the robot actuation is designed as 3 DoFs which are independent from each other. The initialization of can be finished by successively actuating the 3 DoFs in turn with an incremental amount , (i.e., the actuation commands are , , and in turn), and measuring the corresponding displacements , . The initial Jacobian matrix could be constructed as follows:where , . The Jacobian matrix is updated by quadratic programming as follows:
They also proposed to use force sensors to measure the tension of each tendon, therefore optimizing the actuation command with the minimal change. Besides the optimization-based construction of inverse kinematic mapping, machine learning–based methods have also been employed in various works, which are summarized in the following section.
2.2 Methods Utilizing Adaptive Kalman Filter
3 Supervised Learning of Inverse Statics/Kinematics
This section will summarize the methods to learn the desired mapping in control offline, that is, using learning-based algorithms to approximate the statics/kinematics of the entire mapping from the actuation space to the task space. Here, the difference between statics models and kinematics models is briefly explained. A statics model depicts the robot configuration by assuming all forces on the manipulator at rest under equilibrium conditions. A kinematics model describes the robot motion only based on the geometric relationship, without considering the applied forces. For example, in the case of soft cable-driven manipulators, the direct forward statics model maps the cable tensions onto the tip position, while the inverse statics model calculates the cable tensions in order to make the tip on the desired position (
TABLE 1
| Literature | Classification criteria | |||||||
|---|---|---|---|---|---|---|---|---|
| Model | Robot structure | Actuation (length) | Act. & TaskDoFs | Mapping to be learned | Samples | Task | Accuracy/mm (mean/STD/max) | |
| Giorelli et al. | 1-hidden-layer FNN | |||||||
| a. | a. 21 neurons | Silicone conical robot | Tendon (310 mm/280 mm) | a. 2 & 2 | 500 (8:2) | Simulation: path generation. Discrete points | 2.27/1.70/9.30 | |
| b. | b. 34 neurons | b. 3 & 3 | 500 (8:2) | 4.2/2.8/12.3 | ||||
| c. | c. 6 neurons | c. 2 & 2 | 405 (8:2) | 22.88/11.80/59.79 | ||||
| d. | d. 28 neurons | d. 3 & 3 | 395 (8:2) | 7.35/--/22.22 | ||||
| Forward: FNN. Inverse: 1-hidden-layer NN in DSL | CBHA | Pneumatic (NA) | a3 × 2 & 3 | Forward: | 4,096 (7:1.5:1.5) | Comparison: robot and model postures (under the same actuation) | Inverse: 1.1 e−4 (MLP) ∼ 4.1 e−4 (RBF)/--/-- | |
| Inverse: | ||||||||
| Thuruthel et al. | 1-hidden-layer NN. | |||||||
| a. Thuruthel et al. (2016a) | a. 20 neurons | a. BHA | a. Pneumatic (0.9 m) | a. 3 × 3 & 3 | a. 10,000 (7:3) | Simulation: continuous path following in terms of position (P)/orientation (O) | p: <1/1.504/-- | |
| b. Thuruthel et al. (2016b) | b. 40 neurons | b. Silicone conical | b. Tendon (31 cm) | b. 12 & 6 | b. 14,000 (8:2) | p: 8.5/2.8/-- O/°: 3.21/1.71/-- | ||
| ELM, GMR, or KNNR | Silicone serpentine | Tendon (NA) | a2 & 2 | 20,000 | Simulation & real robot: path following | 2.1275∼ 2.5556/−/− | ||
| LWPR online update | Silicone cylindrical | Pneumatic (93 mm) | 3 & 2 | >1,000 for initialization [FEA | Angular path following (2D) + external forces | Free space/°: 0.90/0.65/2.80 | ||
| Disturbed/°: 2.49/1.74/11.03 | ||||||||
| LWPR online update | Silicone cylindrical | Pneumatic (155 mm) | 3 × 2 & 3 | — | Path following (3D) + tip load (72% robot mass) | With load: 0.98/0.26/-- | ||
| LGPR online update | Silicone cylindrical | Pneumatic (67 mm) | 3 & 2 | 300 | Path following (2D visual servo) + tip load | Free space/pixel: 5.4/--/11.5 | ||
| a. 2-hidden-layer FNN (20 × 2) | Silicone cylindrical | Tendon (20 cm) | a. 3 & 2 | a. 308 (9:1) | a. Real robot | Real robot: 6.2∼9.2/−/− | ||
| b: 3-hidden-layer FNN (25 × 3) | b. 3 × 2 & 2 | b. 15414 (8:1:1) | b. Simulation 2D path following | — | ||||
Sampling of inverse statics/kinematics learning-based control for continuum robots. Intended to be exemplary, not comprehensive.
Only the actuation dimensions that are related to the end-effector control are considered; “×2” or “×3” means the number of segments in the manipulator.
(F)NN, (feed-forward) neural network; GMR, Gaussian mixture regression; FEA, finite element analysis; DSL, distal supervised learning; KNNR, K-nearest neighbors regression; STD, standard deviation; (C)BHA, (compact) bionic handling assistant; LWPR, locally weighted projection regression; MLP, multilayer perceptron; ELM, extreme learning machine; LGPR, locally Gaussian process regression; RBF, radial basis function.
3.1 Mapping to Be Learned
As in Figure 2, successive mappings between the actuation and configuration spaces, the configuration and task spaces, or directly from the actuation space to the task space have to be defined as the forward kinematics. The actuator input (at equilibrium) is represented as at time step , where denotes the -dimensional actuation space. Let be the manipulator configuration parameters under input , which corresponds to a specific task space status such as the end-effector position and the orientation normal in the Cartesian space. The collective pose variable can be thus denoted. It should be noticed that for continuum robots, not only can the pose of an appointed point be defined as the control objective but also the entire manipulator shape can be represented by more DoFs. Here, we take the three-dimensional (3D) position as an example to show the representation of the absolute forward robot-independent mapping as follows:and that of absolute forward statics as follows:With quasi-static movements, the forward transition model can be expressed in the incremental format as follows:where is the difference of inputs between time step and , and denotes the incremental displacement. The control objective of inverse statics is to generate an actuation command , thus steering the manipulator to the desired in the task space as follows:
In inverse kinematics for closed-loop control, the change of actuation command , thus achieving the desired movement in the task space, is calculated. Therefore, mapping Eq. 11 is deduced to approximate the inverse kinematics of Eq. 9, as follows:
The inverse transition heavily depends on the last robot configuration that is supposed to be unknown during the training of an operation. However, since during quasi-static movements, the robot configuration can be represented/defined by the corresponding actuation , the inverse kinematics function Eq. 11 can be approximated as follows:here, as long as the sensory information of the task space variable and the encoded actuation command are available, the inverse kinematics can be learned to accomplish various control tasks, without the need of analytical/quantitative modeling. According to the control object, the task space can be specified as the 2D/3D position , the 2D/3D orientation , or be defined in other coordinate frames. For example, in the visual-servoing tasks (
3.2 Learning Approaches
As can be seen from Table 1, neural networks (NNs) are the most commonly used regression model to approximate the mapping. The feed-forward NN (FNN) is the basic type, where the information always flows from the input side to the output side with the weighted calculation of hidden layers (Svozil et al., 1997). If not specified, an NN usually indicates an FNN. Specific types of FNNs like the extreme learning machine (ELM) were also applied. Besides, some regression methods perform satisfying results in continuum robot control, and representative ones can be the locally weighted projection regression (LWPR) (
3.3 Problems to Be Considered
3.3.1 Data Exploration
For the offline trained models, collection and selection of the training data are crucial for the accuracy of the model. Requirements of the samples are as follows: 1) covering the whole workspace of the robot end-effector and 2) evenly distributed in the task space to ensure consistent estimation performance in all areas. Most existing works used a kind of motor babbling approach (Thuruthel et al., 2016b), applying the interpolation in the actuation or configuration space. The number of optimal samples will be dependent on the workspace range and motion step size. Sample (input–output) pairs were collected by incrementally actuating the motors (or other specific mechanisms) with a fixed increment and saving the corresponding end-effector status. With the actuation command in the safe range, this procedure will result in an ergodic dataset. De-noising and filtering were usually needed to abstract high-quality and nonrepetitive samples. To fully exploit the advantage of learning-based approaches, normalizations of the input and output data (into [−1,1] or [0,1]) were conducted before training. For the training of relative mappings, an additional process will be finding out all possible displacements between two random points and filtering out motions in a fixed range of step size (
3.3.2 Structural Optimization of Neural Networks
Although no prior knowledge of the robot modeling is required in data-driven approaches, there are hyper-parameters to be tuned, especially in the NNs, since LWPR and LGPR would not require the manual tuning of any hyper-parameters. The structural optimization of NNs focused on the number of hidden layers and neurons once a specific model was selected. Previous works of
3.3.3 Actuation/Configuration Redundancy
In the learning-based kinematic control, redundancy is the feature which is possible to generate inconsistent samples with the same effector pose but different joint angles or actuation commands. Learning from such examples will lead to invalid solutions (
3.4 Combination of Analytical Model and Learning-Based Component
Either in general or specific tasks, there are hybrid controllers proposed, combining the analytical dynamics/kinematics model and learning-based approaches (e.g., NN) to accomplish robust control performance. Methods combining the learning-based and conventional components also appeared in several works. For example,
4 Reinforcement Learning Strategies
With the development of artificial intelligence, reinforcement learning is emerging in the robotics community, which is a natural application for learning-based control since the interaction between robots and the environment is necessary. Reinforcement learning offers appealing tools enabling to complete sophisticated tasks and accommodate complex environments, which may be limited in conventional control strategies.
Reinforcement learning is regarded as a Markov decision process (MDP), represented using a tuple . In the agent’s interaction with the environment, is the set of the agent’s possible states, where is the current state and is the next state after the agent transition. presents the set of the agent’s actions, where is the action. is the state-transition probability of the agent transiting from the current state to the future state after the implementation of action . The states and actions constitute the trajectory . defines the reward function after the agent executes action at state, and for convenience, represents the immediate reward of one transition and is the accumulated reward or expected return of the whole trajectory, as written in Eq. 13.where is the number of time steps in the trajectory . For the infinite-horizon reinforcement learning problem, the effect of a future reward on the present decision could be considered with the reward discount factor ranging from 0 to 1, which is common for classical reinforcement learning (
For the finite-horizon reinforcement learning problem, the average reward function is considered as shown in Eq. 15.
Policy is the mapping from the state to the action , namely, given the current state, it could suggest the next step to obtain an optimal reward. The value function could evaluate the quality of the policy, offering the quantitative metric for the behavior decision maker. One of the value functions is called the state–value function , which defines the value of state under the policy .
Another one is the action–value function , which could assess the action at state under the policy .
Using Eq. 13 and Eq. 14 and the Bellman equation (
4.1 The Goal of Reinforcement Learning
In the context of mathematics, the goal of reinforcement learning is to explore an optimal policy that could instruct actions based on the present observation. The objective is to maximize the accumulated reward (Eq. 13), which determines the learning task. When considered in robot control, the goal of reinforcement learning is to figure out a control strategy that could generate optimal instruction for robot action in order to accomplish the specified task effectively. The reward function is designed manually to train the robot with certain characteristics, for example, penalizing the times of transition to enable the robot to reach the target in as few movements as possible.
For instance, there is a soft planar robot planned to touch a designated point in 2D space, where reinforcement learning can be explained as below. Sensors on the robot provide the observation about its relative position to the target, as well as moving velocity and direction, which describe the current state . The soft robot is actuated using several inflating air chambers so that it could elongate or contract, thus steering the robot to the left and the right, indicating the action set . Reward function is designed manually, for example, the reward on short relative distance and the penalty on transition times could accelerate the learning convergence process. Policy gives the action suggestion based on the observation of the current state to maximize the cumulative reward. Figure 3 describes the pipeline of reinforcement learning algorithms in soft manipulator control. In the training stage (Figure 3A), trajectories calculated from the forward kinematic model or the simulation environment contribute to training the control policy. In the application stage (Figure 3B), every time upon receiving the sensor-observed state and target information, the learned control policy would give instructed actions, which would be executed by the actuator. Specificities of the application in continuum robots will be introduced in the following section.
FIGURE 3

Schematics of reinforcement learning in soft robot manipulation, with (A) the policy training stage and (B) the application stage separately shown.
4.2 Reinforcement Learning in Soft Robot Manipulation
Compared with the conventional joint-linked robots, the application of reinforcement learning in soft robots may face a number of challenges requiring specific attention. Reinforcement learning enables robots to learn from experience, which demands thousands of interactions with the environment. In addition to the tedious data collection, the noisy data, incomplete observation in practice, and the highly frequent movement could even damage the soft robot since it is mainly actuated hydraulically or pneumatically. To improve the effectiveness of training data collection, the model-free reinforcement learning approach has arisen, which obtains the learning experience from simulations. However, when transferred from the simulation to the prototype, the model-free algorithm might perform with large deviations since the modeling without real data could be inaccurate. Besides, the compliance and flexibility of soft robots cause high dimensional and continuous space, as to which the appropriate state and space discretization methods are expected in reinforcement learning. Furthermore, it is noticed that the soft manipulator would experience attenuation when exerted by external loads or disturbances. Not only could it impede the modeling of robots but it also makes the loading robustness experiment a necessity for the validation of reinforcement learning.
Referring to various criteria, reinforcement learning could be classified into different categories, such as model-based/model-free, policy-based/value-based, and averaged/discounted return function algorithms. The subsequent sections will introduce the first two kinds of categories in detail.
4.2.1 Model-Based Vs. Model-Free Reinforcement Learning
Upon the employment of explicit transition functions, reinforcement learning could be classified into two categories: model-based and model-free algorithms. In robotic applications, model-based methods mostly concentrate on forward kinematics/dynamics models which require prior knowledge on robots or environments (
Model-free reinforcement learning would exploit the virtual training data for policy learning. That is, the data are obtained via robot simulations. It is worth noting that the modeling of the robot and the environment in the simulation is based on simplified assumptions such as PCC (Webster and Jones, 2010;
4.2.1.1 Reinforcement Learning With Kinematics/Dynamics Model
The policy trained on kinematics/dynamics model-based methods can perform stably, while in model-free methods, the derivation may be large when the algorithm is implemented from the simulation into the real robot. This is reasonable as the data from physical interaction are more realistic and persuasive than the one from simulation, where the virtual model is established using many simplifications and assumptions. Moreover, when the robot interacts with the environment in a circumstance that never appeared before, the policy might be invalid.
Thuruthel et al. (2018) leveraged the model-based policy learning algorithm on a simulated tendon-driven soft manipulator capable of touching dynamic targets. A nonlinear autoregressive network with exogenous inputs (NARX) was employed to establish the forward dynamic models using the observed data. Policy iteration from the sampled trajectories was used to give the optimal action directly. Wu et al. (2020) accomplished the position control of a cable-driven soft arm employing Deep Q-learning. Similar to the procedure in the study by Thuruthel et al. (2018), experiment data was collected to model the manipulator in simulation.
However, the shortcoming of kinematics/dynamics model-based reinforcement learning is that the physical interactions would be time-consuming and the data may be noisy; they might even bring more mechanical wear on the robot prototype, particularly the vulnerable soft robot (
4.2.1.2 Reinforcement Learning Without Kinematics/Dynamics Model
In recent years, kinematics/dynamics model-free reinforcement learning implemented on the physical robot has attracted interest in the soft robot community, since it could circumvent the accuracy requirement of analytical modeling, which is hampered by the intrinsic nonlinearity and uncertain external disturbances. In general, the trajectory data are generated in the simulation environment for model-free reinforcement learning, and subsequently, the trained model would be transferred to the physical robot.
You et al. (2017) investigated a model-free reinforcement learning control strategy for a multi-segment soft manipulator, called honeycomb pneumatic networks (HPNs), capable of physically reaching the target in 2D space. The control policy was learned using Q-learning with the simulation data, demonstrating its effectiveness and robustness in simulation and practice.
TABLE 2
| Literature | Reinforcement learning algorithm | Task | Training steps/time | Distance error/success rate |
|---|---|---|---|---|
| Thuruthel et al. (2018) | Policy search | 3D position reaching | 8,000 s | Without load: 0.009∼0.017 m |
| With load: 0.022 m | ||||
| Wu et al. (2020) | Q-learning | 2D position reaching | 1,000 iterations | Without load: <0.5 cm |
| With load: <1 cm | ||||
| You et al. (2017) | Q-learning | 2D position reaching | 1,000 iterations | <10 mm |
| Q-learning | Interaction tasks including drawer opening and handwheel rotating | 120 iterations (about 60 s) and 20,000 iterations (about 11 h) with/without the method of virtual goals | Task success rate 98.86% | |
| DQN | 3D position reaching | 5,000 episodesa | 3.05 cm | |
| You et al. (2019) | DDQN | 3D position reaching | 100 episodes | 6.58 ± 5.6 mm |
| Actor–critic | 3D position control | 300 episodes | — | |
| DDPG | 3D path tracking | 10,000 episodes | ≤3 cm | |
| PPO | 2D tracking with changing goals | 6,400 episodes | — |
Sampling of reinforcement learning control for continuum robots, including the algorithm, task, and training duration. Intended to be exemplary, not comprehensive.
One episode in reinforcement learning means a sequence of states, actions, and rewards, which ends with the terminal state. The time length of one episode depends on the specific task.
4.2.2 Policy-Based Vs. Value-Based Reinforcement Learning
In addition to the perspective of modeling, reinforcement learning algorithms could also be categorized in terms of the solution of optimal policy: the policy search, value-based, and actor–critic methods (
5 Discussion and Conclusion
In this article, we surveyed the state-of-the-art machine learning–based control strategies of continuum robots. Compared with conventional modeling, learning-based mappings provide effective substitutes for analytical models in the feed-forward control loop, without the need of manual model construction and calibration (note: a brief summary of their comparisons can be found in Table 3). The correction load utilizing the feedback loop is therefore reduced. The format of learning objects also varies a lot; in addition to the direct forward/inverse kinematics/statics relationships, learning-based components also contribute to the optimization or compensation units in several works. For example, in the metaheuristics-assisted approach presented in the study by
TABLE 3
| Aspects | Analytical modeling | Supervised learning* | Reinforcement learning | |
|---|---|---|---|---|
| Human Intervention (Single robot) | Model derivation | ✓ | ✗ | ✗ |
| Parameter tuning | ✓ | ✗ | ✗ | |
| Data collection | ✗ | ✓ (offline) | ✓ (online) | |
| Training | ✗ | ✓ | ✓ | |
| Generalization (Same prototype) | Model derivation | ✗ | ✗ | ✗ |
| Parameter tuning | ✓ | ✗ | ✗ | |
| Data collection | ✗ | ✓ (offline) | ✓ (online) | |
| Training | ✗ | ✓ | ✓ | |
| Dependence on data | Low | high | high | |
| Online refinement | ✗ | ✓ | ✓ | |
| Adaptability to un-modeled disturbances | ✗ | ✓ | ✓ | |
A conclusive comparison of analytical-modeling-based and machine-learning-based control.
Cells marked with gray shading indicate advantages.
All the methods mentioned in Sections 3 and 4 facilitate the release of sensor variety demand. However, for those approaches that solely rely on sensory data, a drawback will be the high requirement on the quality of feedback information (i.e., task space sensing). No matter for iterative- or regression-based machine learning techniques, the distribution and accuracy of sensory data would play an important role, which has been discussed in Section 3.3. To adapt to the unstructured external conditions (e.g., contact forces), online updates should be an emphasis in future applications. This is also one major advantage of learning-based control, while relevant attempts have been made in recent works (
Section 5 concludes the various deployments of reinforcement learning in soft continuum robots and reveals its prospect in dealing with complex learning tasks automatically. Nevertheless, it is observed that most previous works concentrated on simplified tasks such as trajectory tracking and goal reaching, which only take little advantage of the powerful learning tool. Developing the learning ability in sophisticated applications is the main challenge of reinforcement learning in soft robots, even the whole robotics field. Additionally, the effectiveness of interaction data collection also hinders further development. Provided with kinematics/dynamics models, enhancing the validity and efficacy of data collection during environmental interactions will be a significant contribution to reinforcement learning. Without such models, the deviation between simulated and actual manipulators would be a big obstacle for reinforcement learning's application. To handle these cases, Sim-to-Real transfer approaches (Zhao et al., 2020) such as domain randomization and domain adaptation may drive a new research focus in the recent future.
Statements
Author contributions
Conceived the format of the review: XW and K-WK; curated the existing bibliography, contributed to figures/materials and the writing of the manuscript: XW and YL; proofreading and revision: XW, YL, and K-WK.
Funding
This study received funding in part from the Research Grants Council (RGC) of Hong Kong (17206818, 17205919, 17207020, and T42-409/18-R), in part from the Innovation and Technology Commission (ITC) (MRP/029/20X and UIM/353), and in part from Multi-Scale Medical Robotics Center Limited, InnoHK, under Grant EW01500.
Conflict of interest
Author XW was employed by the company Multi-Scale Medical Robotics Center Limited.
The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AmouriA.MahfoudiC.ZaatriA.LakhalO.MerzoukiR. (2017). A Metaheuristic Approach to Solve Inverse Kinematics of Continuum Manipulators. Proc. Inst. Mech. Eng. J. Syst. Control. Eng.231 (5), 380–394. 10.1177/0959651817700779
2
AnsariY.MantiM.FaloticoE.CianchettiM.LaschiC. (2017). Multiobjective Optimization for Stiffness and Position Control in a Soft Robot Arm Module. IEEE Robotics Automation Lett.3 (1), 108–115. 10.1109/LRA.2017.2734247
3
AnsariY.MantiM.FaloticoE.MollardY.CianchettiM.LaschiC. (2017). Towards the Development of a Soft Manipulator as an Assistive Robot for Personal Care of Elderly People. Int. J. Adv. Robotic Syst.14 (2), 1729881416687132. 10.1177/1729881416687132
4
AntmanS. S. (2005). “Problems in Nonlinear Elasticity,” in Nonlinear Problems of ElasticityNew York City: Springer, 513–584.
5
BernJ. M.SchniderY.BanzetP.KumarN.CorosS. (2020). “Soft Robot Control with a Learned Differentiable Model,” in 2020 3rd IEEE International Conference on Soft Robotics, Yale University, USA, May 15–July 15, 2020 (RoboSoft), 417–423. 10.1109/robosoft48309.2020.9116011
6
BraganzaD.DawsonD. M.WalkerI. D.NathN. (2007). A Neural Network Controller for Continuum Robots. IEEE Trans. Robot.23 (6), 1270–1277. 10.1109/tro.2007.906248
7
Burgner-KahrsJ.RuckerD. C.ChosetH. (2015). Continuum Robots for Medical Applications: A Survey. IEEE Trans. Robot.31 (6), 1261–1280. 10.1109/tro.2015.2489500
8
ChenJ.LauH. Y. (2016). “Learning the Inverse Kinematics of Tendon-Driven Soft Manipulators with K-Nearest Neighbors Regression and Gaussian Mixture Regression,” in 2016 2nd International Conference on Control, Hong Kong, China, Apr 28–30, 2016 (Automation and Robotics (ICCAR), 103–107. 10.1109/iccar.2016.7486707
9
FangG.WangX.WangK.LeeK.-H.HoJ. D. L.FuH.-C.et al (2019). Vision-based Online Learning Kinematic Control for Soft Robots Using Local Gaussian Process Regression. IEEE Robot. Autom. Lett.4 (2), 1194–1201. 10.1109/lra.2019.2893691
10
FuH.-C.HoJ. D. L.LeeK.-H.HuY. C.AuS. K. W.ChoK.-J.et al (2020). Interfacing Soft and Hard: A spring Reinforced Actuator. Soft Robotics7 (1), 44–58. 10.1089/soro.2018.0118
11
George ThuruthelT.AnsariY.FaloticoE.LaschiC. (2018). Control Strategies for Soft Robotic Manipulators: A Survey. Soft robotics5 (2), 149–163. 10.1089/soro.2017.0007
12
GiorelliM.RendaF.CalistiM.ArientiA.FerriG.LaschiC. (2012). “A Two Dimensional Inverse Kinetics Model of a cable Driven Manipulator Inspired by the octopus Arm,” in 2012 IEEE international conference on robotics and automation, Saint Paul, MN, USA, May 14–19, 2012, 3819–3824. 10.1109/icra.2012.6225254
13
GiorelliM.RendaF.CalistiM.ArientiA.FerriG.LaschiC. (2015a). Neural Network and Jacobian Method for Solving the Inverse Statics of a Cable-Driven Soft Arm with Nonconstant Curvature. IEEE Trans. Robot.31 (4), 823–834. 10.1109/tro.2015.2428511
14
GiorelliM.RendaF.CalistiM.ArientiA.FerriG.LaschiC. (2015b). Learning the Inverse Kinetics of an Octopus-Like Manipulator in Three-Dimensional Space. Bioinspir. Biomim.10 (3), 035006. 10.1088/1748-3190/10/3/035006
15
GiorelliM.RendaF.FerriG.LaschiC. (2013a). “A Feed Forward Neural Network for Solving the Inverse Kinetics of Non-constant Curvature Soft Manipulators Driven by Cables,” in Dynamic Systems and Control Conference, Palo Alto, CA, USA, Oct 21–23, 2013 (Stanford University), V003T38A001. 10.1115/dscc2013-3740
16
GiorelliM.RendaF.FerriG.LaschiC. (2013b). “A Feed-Forward Neural Network Learning the Inverse Kinetics of a Soft Cable-Driven Manipulator Moving in Three-Dimensional Space,” in 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, Nov 3–7, 2013, 5033–5039. 10.1109/iros.2013.6697084
17
HoJ. D. L.LeeK.-H.TangW. L.HuiK.-M.AlthoeferK.LamJ.et al (2018). Localized Online Learning-Based Control of a Soft Redundant Manipulator under Variable Loading. Adv. Robotics32 (21), 1168–1183. 10.1080/01691864.2018.1528178
18
HolstenF.Engell-NørregårdM. P.DarknerS.ErlebenK. (2019). “Data Driven Inverse Kinematics of Soft Robots Using Local Models,” in 2019 International Conference on Robotics and Automation (ICRA), Montreal, Canada, May 20–24, 2019, 6251–6257. 10.1109/icra.2019.8794191
19
JiangH.LiuX.ChenX.WangZ.JinY.ChenX. (2016). “Design and simulation analysis of a soft manipulator based on honeycomb pneumatic networks,” in 2016 IEEE International Conference on Robotics and Biomimetics (ROBIO), Qingdao, China, December 3–7, 2016, 350–356. 10.1109/ROBIO.2016.7866347
20
JiangH.WangZ.JinY.ChenX.LiP.GanY.et al (2021). Hierarchical Control of Soft Manipulators towards Unstructured Interactions. Int. J. Robotics Res.40 (1), 411–434. 10.1177/0278364920979367
21
JonesB. A.WalkerI. D. (2006). Kinematics for Multisection Continuum Robots. IEEE Trans. Robot.22 (1), 43–55. 10.1109/tro.2005.861458
22
JonesB. A.WalkerI. D. (2006). Practical Kinematics for Real-Time Implementation of Continuum Robots. IEEE Trans. Robot.22 (6), 1087–1099. 10.1109/tro.2006.886268
23
KangR.BransonD. T.ZhengT.GuglielminoE.CaldwellD. G. (2013). Design, Modeling and Control of a Pneumatically Actuated Manipulator Inspired by Biological Continuum Structures. Bioinspir. Biomim.8 (3), 036008. 10.1088/1748-3182/8/3/036008
24
KoberJ.BagnellJ. A.PetersJ. (2013). Reinforcement Learning in Robotics: A Survey. Int. J. Robotics Res.32 (11), 1238–1274. 10.1177/0278364913495721
25
LeeK.-H.FuD. K. C.LeongM. C. W.ChowM.FuH.-C.AlthoeferK.et al (2017a). Nonparametric Online Learning Control for Soft Continuum Robot: An Enabling Technique for Effective Endoscopic Navigation. Soft robotics4 (4), 324–337. 10.1089/soro.2016.0065
26
LeeK.-H.LeongM. C.ChowM. C.FuH.-C.LukW.SzeK.-Y.YeungC.-K.KwokK.-W. (2017b). “FEM-Based Soft Robotic Control Framework for Intracavitary Navigation,” in 2017 IEEE International Conference on Real-time Computing and Robotics (RCAR), Okinawa, Japan, Jul 14–18, 2017, 11–16. 10.1109/rcar.2017.8311828
27
LeeK.-H.FuK. C. D.GuoZ.DongZ.LeongM. C. W.CheungC.-L.et al (2018). MR Safe Robotic Manipulator for MRI-Guided Intracardiac Catheterization. Ieee/asme Trans. Mechatron.23 (2), 586–595. 10.1109/tmech.2018.2801787
28
LiM.KangR.BransonD. T.DaiJ. S. (2017). Model-Free Control for Continuum Robots Based on an Adaptive Kalman Filter. IEEE/ASME Trans. Mechatronics23 (1), 286–297. 10.1109/TMECH.2017.2775663
29
LiuX.GasotoR.JiangZ.OnalC.FuJ. (2020). “Learning to Locomote with Artificial Neural-Network and Cpg-Based Control in a Soft Snake Robot,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, USA, Oct 25–29, 2020, 7758–7765. 10.1109/iros45743.2020.9340763
30
LunT. L. T.WangK.HoJ. D. L.LeeK.-H.SzeK. Y.KwokK.-W. (2019). Real-Time Surface Shape Sensing for Soft and Flexible Structures Using Fiber Bragg Gratings. IEEE Robot. Autom. Lett.4 (2), 1454–1461. 10.1109/lra.2019.2893036
31
LunzeJ. (1998). Qualitative Modelling of Dynamical Systems: Motivation, Methods, and Prospective Applications. Mathematics Comput. Simulation46 (5-6), 465–483. 10.1016/s0378-4754(98)00077-9
32
MahlerJ.KrishnanS.LaskeyM.SenS.MuraliA.KehoeB.PatilS.WangJ.FranklinM.AbbeelP. (2014). “Learning Accurate Kinematic Control of Cable-Driven Surgical Robots Using Data Cleaning and Gaussian Process Regression,” in 2014 IEEE international conference on automation science and engineering (CASE), Taipei, Taiwan, Aug 18–22, 2014, 532–539. 10.1109/coase.2014.6899377
33
MelinguiA.EscandeC.BenoudjitN.MerzoukiR.MbedeJ. B. (2014). Qualitative Approach for Forward Kinematic Modeling of a Compact Bionic Handling Assistant Trunk. IFAC Proc. Volumes47 (3), 9353–9358. 10.3182/20140824-6-za-1003.01758
34
MelinguiA.LakhalO.DaachiB.MbedeJ. B.MerzoukiR. (2015). Adaptive Neural Network Control of a Compact Bionic Handling Arm. Ieee/asme Trans. Mechatron.20 (6), 2862–2875. 10.1109/tmech.2015.2396114
35
MelinguiA.MerzoukiR.MbedeJ. B.EscandeC.DaachiB.BenoudjitN. (2014). “Qualitative Approach for Inverse Kinematic Modeling of a Compact Bionic Handling Assistant Trunk,” in 2014 International Joint Conference on Neural Networks (IJCNN), Beijing, China, Jul 6–11, 2014, 754–761. 10.1109/ijcnn.2014.6889947
36
MoerlandT. M.BroekensJ.JonkerC. M. (2020). Model-Based Reinforcement Learning: A Survey. arXiv preprint arXiv:.16712.
37
NajarA.ChetouaniM. (2020). Reinforcement Learning with Human Advice. A Survey. arXiv preprint arXiv:.11016
38
NordmannA.RolfM.WredeS. (2012). “Software Abstractions for Simulation and Control of a Continuum Robot,” in International Conference on Simulation, Modeling, and Programming for Autonomous Robots, Tsukuba, Japan, Nov 5–8, 2012, 113–124. 10.1007/978-3-642-34327-8_13
39
PetersJ.SchaalS. (2008). Learning to Control in Operational Space. Int. J. Robotics Res.27 (2), 197–212. 10.1177/0278364907087548
40
PolydorosA. S.NalpantidisL. (2017). Survey of Model-Based Reinforcement Learning: Applications on Robotics. J. Intell. Robot Syst.86 (2), 153–173. 10.1007/s10846-017-0468-y
41
QueißerJ. F.NeumannK.RolfM.ReinhartR. F.SteilJ. J. (2014). “An Active Compliant Control Mode for Interaction with a Pneumatic Soft Robot,” in 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA, Sep 14–18, 2014, 573–579.
42
ReinhartR. F.SteilJ. J. (2016). Hybrid Mechanical and Data-Driven Modeling Improves Inverse Kinematic Control of a Soft Robot. Proced. Technol.26, 12–19. 10.1016/j.protcy.2016.08.003
43
ReinhartR.ShareefZ.SteilJ. (2017). Hybrid Analytical and Data-Driven Modeling for Feed-Forward Robot Control †. Sensors17 (2), 311. 10.3390/s17020311
44
RobinsonG.DaviesJ. B. C. (1999). “Continuum Robots-A State of the Art,” in Proceedings 1999 IEEE international conference on robotics and automation (Cat. No. 99CH36288C), Detroit, MI, USA, May 10–15, 1999, 2849–2854.
45
RolfM.SteilJ. J. (2013). Efficient Exploratory Learning of Inverse Kinematics on a Bionic Elephant Trunk. IEEE Trans. Neural networks Learn. Syst.25 (6), 1147–1160. 10.1109/TNNLS.2013.2287890
46
RolfM.SteilJ. J.GiengerM. (2010). Goal Babbling Permits Direct Learning of Inverse Kinematics. IEEE Trans. Auton. Ment. Dev.2 (3), 216–229. 10.1109/tamd.2010.2062511
47
Rong-HuC.Zhong-ShengH. (2007). Dual-Stage Optimal Iterative Learning Control for Nonlinear Non-Affine Discrete-Time Systems. Acta Automatica Sinica33 (10), 1061–1065. 10.1360/aas-007-1061
48
RuckerD. C.JonesB. A.Webster IIIR. J.III (2010). A Geometrically Exact Model for Externally Loaded Concentric-Tube Continuum Robots. IEEE Trans. Robot.26 (5), 769–780. 10.1109/tro.2010.2062570
49
SatheeshbabuS.UppalapatiN. K.ChowdharyG.KrishnanG. (2019). “Open Loop Position Control of Soft Continuum Arm Using Deep Reinforcement Learning,” in 2019 International Conference on Robotics and Automation (ICRA), Montreal, Canada, May 20–24, 2019, 5133–5139. 10.1109/icra.2019.8793653
50
SatheeshbabuS.UppalapatiN. K.FuT.KrishnanG. (2020). “Continuous Control of a Soft Continuum Arm Using Deep Reinforcement Learning,” in 2020 3rd IEEE International Conference on Soft Robotics (RoboSoft), United States, May 15–July 15, 2020 (Yale University), 497–503.
51
SicilianoB.KhatibO. (2016). Springer Handbook of Robotics. Berlin: Springer.
52
SubudhiB.MorrisA. S. (2009). Soft Computing Methods Applied to the Control of a Flexible Robot Manipulator. Appl. Soft Comput.9 (1), 149–158. 10.1016/j.asoc.2008.02.004
53
SuttonR. S.BartoA. G. (2018). Reinforcement Learning: An Introduction. Cambridge, MA: MIT press.
54
SvozilD.KvasnickaV.PospichalJ. (1997). Introduction to Multi-Layer Feed-Forward Neural Networks. Chemometrics Intell. Lab. Syst.39 (1), 43–62. 10.1016/s0169-7439(97)00061-0
55
TangZ. Q.HeungH. L.TongK. Y.LiZ. (2019a). “A Novel Iterative Learning Model Predictive Control Method for Soft Bending Actuators,” in 2019 International Conference on Robotics and Automation (ICRA), Montreal, Canada, May 20–24, 2019, 4004–4010. 10.1109/icra.2019.8793871
56
TangZ. Q.HeungH. L.TongK. Y.LiZ. (2019b). Model-Based Online Learning and Adaptive Control for a "Human-Wearable Soft Robot" Integrated System. Int. J. Robotics Res.40 (1), 256–276. 10.1177/0278364919873379
57
ThuruthelT. G.FaloticoE.CianchettiM.LaschiC. (2016a). “Learning Global Inverse Kinematics Solutions for a Continuum Robot,” in Symposium on Robot Design, Lisbon, Portugal, July 29–31, 2016, (Dynamics and Control), 47–54. 10.1007/978-3-319-33714-2_6
58
ThuruthelT.FaloticoE.CianchettiM.RendaF.LaschiC. (2016b). “Learning Global Inverse Statics Solution for a Redundant Soft Robot,” in Proceedings of the 13th International Conference on Informatics in Control, Udine, Italy, Jun 20–23, 2016 (Automation and Robotics), 303–310. 10.5220/0005979403030310
59
ThuruthelT. G.FaloticoE.RendaF.LaschiC. (2018). Model-based Reinforcement Learning for Closed-Loop Dynamic Control of Soft Robotic Manipulators. IEEE Trans. Robotics35 (1), 124–134. 10.1109/TRO.2018.2878318
60
TrivediD.LotfiA.RahnC. D. (2008). Geometrically Exact Models for Soft Robotic Manipulators. IEEE Trans. Robot.24 (4), 773–780. 10.1109/tro.2008.924923
61
WangK.MakC. H.HoJ. D.-L.LiuZ.-Y.SzeK. Y.WongK. K.et al (2021). Large-scale Surface Shape Sensing with Learning-Based Computational Mechanics. Adv. Intell. Syst. 10.1002/aisy.202100089(Accepted)
62
WangX.FangG.WangK.XieX.LeeK.-H.HoJ. D. L.et al (2020). Eye-in-Hand Visual Servoing Enhanced with Sparse Strain Measurement for Soft Continuum Robots. IEEE Robot. Autom. Lett.5 (2), 2161–2168. 10.1109/lra.2020.2969953
63
WangX.LeeK.-H.FuD. K. C.DongZ.WangK.FangG.et al (2018). Experimental Validation of Robot-Assisted Cardiovascular Catheterization: Model-Based versus Model-free Control. Int. J. CARS13 (6), 797–804. 10.1007/s11548-018-1757-z
64
WangZ.SchaulT.HesselM.HasseltH.LanctotM.FreitasN. (2016). “Dueling Network Architectures for Deep Reinforcement Learning,” in International conference on machine learning, New York City, NY, USA, Jun 19–24, 2016, 1995–2003.
65
WebsterR. J.JonesB. A. (2010). Design and Kinematic Modeling of Constant Curvature Continuum Robots: a Review. Int. J. Robotics Res.29 (13), 1661–1683. 10.1177/0278364910368147
66
WuQ.GuY.LiY.ZhangB.ChepinskiyS. A.WangJ.et al (2020). Position Control of Cable-Driven Robotic Soft Arm Based on Deep Reinforcement Learning. Information11 (6), 310. 10.3390/info11060310
67
XuW.ChenJ.LauH. Y. K.RenH. (2017). Data-Driven Methods towards Learning the Highly Nonlinear Inverse Kinematics of Tendon-Driven Surgical Manipulators. Int. J. Med. Robotics Comput. Assist. Surg.13 (3), e1774. 10.1002/rcs.1774
68
YipM. C.CamarilloD. B. (2014). Model-less Feedback Control of Continuum Manipulators in Constrained Environments. IEEE Trans. Robot.30 (4), 880–889. 10.1109/tro.2014.2309194
69
YipM. C.CamarilloD. B. (2016). Model-less Hybrid Position/force Control: a Minimalist Approach for Continuum Manipulators in Unknown, Constrained Environments. IEEE Robot. Autom. Lett.1 (2), 844–851. 10.1109/lra.2016.2526062
70
YipM. C.SgangaJ. A.CamarilloD. B. (2017). Autonomous Control of Continuum Robot Manipulators for Complex Cardiac Ablation Tasks. J. Med. Robot. Res.02 (01), 1750002. 10.1142/s2424905x17500027
71
YouH.BaeE.MoonY.KweonJ.ChoiJ. (2019). Automatic Control of Cardiac Ablation Catheter with Deep Reinforcement Learning Method. J. Mech. Sci. Technol.33 (11), 5415–5423. 10.1007/s12206-019-1036-0
72
YouX.ZhangY.ChenX.LiuX.WangZ.JiangH.ChenX. (2017). “Model-free Control for Soft Manipulators Based on Reinforcement Learning,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, Canada, Sep 24–28, 2017, 2909–2915. 10.1109/iros.2017.8206123
73
ZhaoWQueraltaJ PWesterlundT (2020). “Sim-to-real transfer in deep reinforcement learning for robotics: a survey,” in in IEEE Symposium Series on Computational Intelligence (SSCI), Canberra, Australia, Vancouver, Canada, December 1–4, 2020, 1737–744. 10.1109/SSCI47803.2020.9308468
Summary
Keywords
continuum robots, data-driven control, inverse kinematics (IK), kinematic/dynamic model-free control, learning-based control, machine learning, reinforcement learning, soft robots
Citation
Wang X, Li Y and Kwok K-W (2021) A Survey for Machine Learning-Based Control of Continuum Robots. Front. Robot. AI 8:730330. doi: 10.3389/frobt.2021.730330
Received
24 June 2021
Accepted
17 August 2021
Published
24 September 2021
Volume
8 - 2021
Edited by
Noman Naseer, Air University, Pakistan
Reviewed by
Deepak Trivedi, General Electric, United States
Muhammad Jawad Khan, National University of Sciences and Technology (NUST), Pakistan
Updates

Check for updates
Copyright
© 2021 Wang, Li and Kwok.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Ka-Wai Kwok, kwokkw@hku.hk
This article was submitted to Biomedical Robotics, a section of the journal Frontiers in Robotics and AI
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.