### Abstract

We consider optimization problems in which the objective requires an inner loop with many steps or is the limit of a sequence of increasingly costly approximations. Meta-learning, training recurrent neural networks, and optimization of the solutions to differential equations are all examples of optimization problems with this character. In such problems, it can be expensive to compute the objective function value and its gradient, but truncating the loop or using less accurate approximations can induce biases that damage the overall solution. We propose randomized telescope (RT) gradient estimators, which represent the objective as the sum of a telescoping series and sample linear combinations of terms to provide cheap unbiased gradient estimates. We identify conditions under which RT estimators achieve optimization convergence rates independent of the length of the loop or the required accuracy of the approximation. We also derive a method for tuning RT estimators online to maximize a lower bound on the expected decrease in loss per unit of computation. We evaluate our adaptive RT estimators on a range of applications including meta-optimization of learning rates, variational inference of ODE parameters, and training an LSTM to model long sequences.

Original language | English (US) |
---|---|

Title of host publication | 36th International Conference on Machine Learning, ICML 2019 |

Publisher | International Machine Learning Society (IMLS) |

Pages | 836-854 |

Number of pages | 19 |

ISBN (Electronic) | 9781510886988 |

State | Published - Jan 1 2019 |

Event | 36th International Conference on Machine Learning, ICML 2019 - Long Beach, United States Duration: Jun 9 2019 → Jun 15 2019 |

### Publication series

Name | 36th International Conference on Machine Learning, ICML 2019 |
---|---|

Volume | 2019-June |

### Conference

Conference | 36th International Conference on Machine Learning, ICML 2019 |
---|---|

Country | United States |

City | Long Beach |

Period | 6/9/19 → 6/15/19 |

### All Science Journal Classification (ASJC) codes

- Education
- Computer Science Applications
- Human-Computer Interaction

## Fingerprint Dive into the research topics of 'Efficient optimization of loops and limits with randomized telescoping sums'. Together they form a unique fingerprint.

## Cite this

*36th International Conference on Machine Learning, ICML 2019*(pp. 836-854). (36th International Conference on Machine Learning, ICML 2019; Vol. 2019-June). International Machine Learning Society (IMLS).