Skip to main navigation Skip to search Skip to main content

Orthogonal Recurrent Neural Networks with Scaled Cayley Transform

  • Kyle E. Helfrich
  • , Devin Willmott
  • , Qiang Ye

Research output: Contribution to journalConference articlepeer-review

68 Scopus citations

Abstract

Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simpler and novel update scheme to maintain orthogonal recurrent weight matrices without using complex valued matrices. This is done by parametrizing with a skew-symmetric matrix using the Cayley transform; such a parametrization is unable to represent matrices with negative one eigenvalues, but this limitation is overcome by scaling the recurrent weight matrix by a diagonal matrix consisting of ones and negative ones. The proposed training scheme involves a straightforward gradient calculation and update step. In several experiments, the proposed scaled Cayley orthogonal recurrent neural network (scoRNN) achieves superior results with fewer trainable parameters than other unitary RNNs.

Original languageEnglish
Pages (from-to)1969-1978
Number of pages10
JournalProceedings of Machine Learning Research
Volume80
StatePublished - 2018
Event35th International Conference on Machine Learning, ICML 2018 - Stockholm, Sweden
Duration: Jul 10 2018Jul 15 2018

Bibliographical note

Publisher Copyright:
© 2018 by the author(s).

Funding

This research was supported in part by NSF Grants DMS-1317424 and DMS-1620082.

FundersFunder number
National Science Foundation Arctic Social Science ProgramDMS-1317424, DMS-1620082

    ASJC Scopus subject areas

    • Software
    • Control and Systems Engineering
    • Statistics and Probability
    • Artificial Intelligence

    Fingerprint

    Dive into the research topics of 'Orthogonal Recurrent Neural Networks with Scaled Cayley Transform'. Together they form a unique fingerprint.

    Cite this