The tokenizer called at the input_ids of example 2 is currently encoding text_1. I think this should be changed to text_2.