Studies / Research

LSTM and Transformer Language Models

By Žan Pušenjak

Deep Learning, FRI, University of Ljubljana


deep-learning

Introduction

The goal of this project was to implement the LSTM and Transformer models and use them on a series of text generating tasks with differentiating training parameters and compare the training and testing results.

Implementation

We used the LSTM PyTorch implementation1 and Attention is all you need2 paper to implement the logic for the models.

We than treined the models on a dataset of Shakespeares work3

Training results

We trained the models on sequence lengths of 32, 64, and 128, and evaluated both training efficiency and output quality.

LSTM model

As expected, shorter sequence lengths resulted in faster training due to reduced model complexity. However, this came at a cost in output quality.

Comparing the generated samples (see Examples [lst:32a][lst:128b]), shorter sequence models produce less coherent text. They tend to form incomplete sentences, lose structure quickly, and fall into repetitive loops. In contrast, longer sequence models generate more structured and contextually consistent text, although repetition is still present.

O Romeo, wherefore art thou doth the state and stand the state and some stand to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove to the prove
O Romeo, wherefore art thous

I mor the the tist wherd wert mess ang mon thet ar tee thas, the ther thin mes shese
I mente atee mhe me wo the sin sone her mo tho s sounse wo mo hithes anote s ee thithis, bo lo het an ar thord the maneethone she to s seet wan toe t it so she thet sore he thate won mane, warde this witiere
houl
O Romeo, wherefore art thou art the consul to the world
With the son and the state of the sease the sease
The senter of the senter of the sease the sease
The senter of the senter of the sease the sease
The senter of the senter of the sease the sease
The senter of the senter of the sease the sease
The senter of the senter of t
O Romeo, wherefore art thou art
the prepitions to the wond on a known heart
Had sound the matcher of my honour, what
I had night of his friends the sentents
Would that will be a lood to be moness to be a parish.
House, that thou with my said and thy sord weed and state.

Lord Mespora:
You soon arms, to bane the laster him.

Transformer model

The Transformer models produced better results overall, particularly in maintaining sentence structure and readability. However, the same pattern holds: shorter sequence lengths train faster but generate lower-quality text.

Compared to the LSTM outputs, Transformer-generated text shows improved coherence and longer-range dependencies, especially at sequence length 128 (see Examples [tr:32] and [tr:128]). Still, repetition and token-level noise remain visible.

Here's to my love! O true apothecary! Thy drugs are quick. e t thonitheaghthethuest teleare bresed ainges.edousckessseat


Ho walsthe s ane, med t waschack wanarene thellesthe inonethoncestheathe
Asckellle our s, tongenet we t iteadeareneeno tourevowo moun ouenched
Thencke tere weshence antheque heelenthellllit.
Houe, wouren mss;

Thty thare honckisheeeay, hemealougele hest a in thengantes,
Heneveatholllleainge t blenashoury mest ancy, ithoue hed, thiner
Here's to my love! O true apothecary! Thy drugs are quick.

Within thou art, thou must needs with a bank and love,
And ask the rock of that school-mates.
That I am dead man ever
The water-mattter, then at the end.

DUCHESSS OF YORK:
Sweet York, speak, now he might hands
My pity mother than a mewelesss of my steeed--balack thee,
As fondly set debile him, and how he mart
an another's head: and 'twillt be done with him,
And all his pale, where next not wha

Conclusion

Like expected the longer sequences produced better results but took longer to learn. The results of the Transformer models outperformed the LSTM but required more epochs to train and could still be trained further.


  1. https://docs.pytorch.org/docs/stable/generated/torch.nn.LSTM.html↩︎

  2. https://arxiv.org/abs/1706.03762↩︎

  3. https://raw.githubusercontent.com/karpathy/char-rnn/master/data/tinyshakespeare/input.txt.↩︎