Stuck on an issue?

Lightrun Answers was designed to reduce the constant googling that comes with debugging 3rd party libraries. It collects links to all the places you might be looking at while hunting down a tough bug.

And, if you’re still stuck at the end, we’re happy to hop on a call to see how we can help out.

KeyError: 'labels' in distill_classifier.py

See original GitHub issue

Environment info

transformers version: 4.6.1
Platform: Darwin-19.6.0-x86_64-i386-64bit
Python version: 3.7.6
PyTorch version (GPU?): 1.8.1 (False)
Tensorflow version (GPU?): 2.2.0 (False)
Using GPU in script?: No
Using distributed or parallel set-up in script?: No

Issue

I am trying to run the distill_classifier.py script from transformers/examples/research_projects/zero-shot-distillation/ with my own text data set and labels on the roberta-large-mnli model. There are a few hundred rows of text and 13 class labels. I am running the following in a cell of my notebook:

!python transformers/examples/research_projects/zero-shot-distillation/distill_classifier.py \
--data_file ./distill_data/train_unlabeled.txt \
--class_names_file ./distill_data/class_names.txt \
--teacher_name_or_path roberta-large-mnli \
--hypothesis_template "This text is about {}." \
--output_dir ./my_student/distilled

The script starts to run but after a short while I receive the following error:

Trainer is attempting to log a value of "{'Science': 0, 'Math': 1, 'Social Studies': 2, 'Language Arts': 3, 'Statistics': 4, 'Calculus': 5, 'Linear Algebra': 6, 'Probability': 7, 'Chemistry': 8, 'Biology': 9, 'Supply chain management': 10, 'Economics': 11, 'Pottery': 12}" 
for key "label2id" as a parameter. 
MLflow's log_param() only accepts values no longer than 250 characters so we dropped this attribute.
 0%|                                                     | 0/7 [00:00<?, ?it/s]Traceback (most recent call last):
  File "transformers/examples/research_projects/zero-shot-distillation/distill_classifier.py", line 338, in <module>
    main()
  File "transformers/examples/research_projects/zero-shot-distillation/distill_classifier.py", line 328, in main
    trainer.train()
  File "/opt/anaconda3/lib/python3.7/site-packages/transformers/trainer.py", line 1272, in train
    tr_loss += self.training_step(model, inputs)
  File "/opt/anaconda3/lib/python3.7/site-packages/transformers/trainer.py", line 1734, in training_step
    loss = self.compute_loss(model, inputs)
  File "transformers/examples/research_projects/zero-shot-distillation/distill_classifier.py", line 119, in compute_loss
    target_p = inputs["labels"]
  File "/opt/anaconda3/lib/python3.7/site-packages/transformers/tokenization_utils_base.py", line 231, in __getitem__
    return self.data[item]
KeyError: 'labels'
  0%|                                                     | 0/7 [00:01<?, ?it/s]

I have re-examined my labels files and am exactly following this guide for distill_classifier.py

https://colab.research.google.com/drive/1mjBjd0cR8G57ZpsnFCS3ngGyo5nCa9ya?usp=sharing&utm_campaign=Hugging%2BFace&utm_medium=web&utm_source=Hugging_Face_8#scrollTo=ECt06ndcnpyb

Any help would be appreciated to distill!

Edit: Updated torch to latest version and receiving the same error. I reduced number of classes from 24 to 13 and still have this issue. When I print inputs in the compute loss function it looks like there is no key for labels:

{'attention_mask': tensor([[1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0]]), 'input_ids': tensor([[ 101, 1999, 2262,  ...,    0,    0,    0],
        [ 101, 4117, 2007,  ...,    0,    0,    0],
        [ 101, 2130, 2295,  ...,    0,    0,    0],
        ...,
        [ 101, 1999, 2760,  ...,    0,    0,    0],
        [ 101, 2057, 6614,  ...,    0,    0,    0],
        [ 101, 2057, 1521,  ...,    0,    0,    0]])}

Is there an additional parameter that is needed to assign the labels?

Edit 2: Just let the colab notebook “Distilling Zero Shot Classification.ipynb” run for a few hours and am receiving the same error with the agnews dataset. It looks like the code in the colab notebook might have an incompatibility with some other files possibly.

Edit 3: I have changed datasets and reduced to 3 classes and tried to add the label_names argument

--label_names ["Carbon emissions", "Energy efficiency", "Water scarcity"]

my ./distill_data/class_names.txt file looks like:

Carbon Emissions
Energy Efficiency 
Water Scarcity

and am still facing the same error.

Who can help

@LysandreJik @sgugger @joeddav

Issue Analytics

State:
Created 2 years ago
Comments:5 (1 by maintainers)

Top GitHub Comments

2reactions

adamFinastracommented, Jun 18, 2021

That did the trick!

1reaction

GaussianNeuronscommented, May 26, 2021

I am facing the same issue and cannot run the google colab examples either. Any help is appreciated!

Top Results From Across the Web

KeyError "labels" occurring in distill_classifier.py official ...

Running the official Distilling Zero Shot Classification.ipynb results in a KeyError: 'labels'. Here is the full output for reference:.

KeyError: 'labels [data] not contained in axis' - Stack Overflow

I have tried several different methods to add a row to an existing Pandas Dataframe. For example I tried the solution here. However...

KeyError Pandas – How To Fix - Data Independent

Pandas KeyError - This annoying error means that Pandas can not find your column name in your dataframe. Here's how to fix this...

How to Fix in Pandas: KeyError: "['Label'] not found in axis"

This tutorial explains how to fix the following error in pandas: KeyError: "['Label'] not found in axis", including an example.

How to Fix: KeyError in Pandas - GeeksforGeeks

In this article, we will discuss how to fix the KeyError in pandas. Pandas KeyError occurs when we try to access some column/row...

Troubleshoot Live Code

Lightrun enables developers to add logs, metrics and snapshots to live code - no restarts or redeploys required.

Start Free

Top Related Reddit Thread

No results found

Top Related Tweet

No results found

Top Related Dev.to Post

No results found

KeyError: 'labels' in distill_classifier.py

Environment info

Issue

Who can help

Issue Analytics

Top GitHub Comments

Top Results From Across the Web

Top Related Medium Post

Top Related StackOverflow Question

Troubleshoot Live Code

Top Related Reddit Thread

Top Related Hackernoon Post

Top Related Tweet

Top Related Dev.to Post

Top Related Hashnode Post

Add a new pipeline for the Relation Extraction task.

Multi-node training for casual language modeling example does not work