NVIDIA/Megatron-LM

[BUG] GPT2BPETokenizer (and possibly others) missing decode and offsets methods

Open

#1,633 opened on Jun 13, 2025

View on GitHub
 (3 comments) (0 reactions) (0 assignees)Python (1,723 forks)batch import
bugcommunity-requestgood first issuemodule: data pipelinewaiting-on-customer

Repository metrics

Stars
 (7,565 stars)
PR merge metrics
 (Avg merge 12d 9h) (245 merged PRs in 30d)

Description

I've encountered an issue when running text generation with Megatron-LM. Apologies in advance if there are any mistakes — I'm a new user.

After successfully preprocessing the data using:

python tools/preprocess_data.py --tokenizer-type GPT2BPETokenizer

and completing pretraining, I tried running text generation using:

tools/run_text_generation_server.py

However, I received the following errors:

AttributeError: '_GPT2BPETokenizer' object has no attribute 'decode'
AttributeError: '_GPT2BPETokenizer' object has no attribute 'offsets'

These errors seem to originate from:

It seems like _GPT2BPETokenizer may be missing decode and offsets methods required by the inference script. Any guidance on how to resolve this would be greatly appreciated.

Contributor guide