Compare commits

...
Author SHA1 Message Date
Nicolas Patry 3bae40495e Checking output format + check raises ValueError (#8986) 2020-12-08 12:25:58 -05:00
LysandreJik 4e244dc1be Docs 2020-11-30 17:59:23 -05:00
LysandreJik eebee0e30a Update documentation 2020-11-30 17:58:40 -05:00
LysandreJik 9a7ae2adc9 return_raw_outputs becomes function_to_apply 2020-11-30 17:39:57 -05:00
Lysandre df2cddfa98 Proposal 2020-11-30 14:25:10 -05:00
Lysandre 6db6c99d1b Switch flag 2020-11-30 14:25:10 -05:00
Lysandre 7833539a9b Return raw outputs in TextClassificationPipeline 2020-11-30 14:25:10 -05:00
LysandreJik 5fd3d81ec9 fix pypi complaint on version naming 2020-11-30 13:54:52 -05:00
Funtowicz MorganandSylvain Gugger 51b071313b Attempt to fix Flax CI error(s) (#8829)
* Slightly increase tolerance between pytorch and flax output

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* test_multiple_sentences doesn't require torch

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Simplify parameterization on "jit" to use boolean rather than str

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Use `require_torch` on `test_multiple_sentences` because we pull the weight from the hub.

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Rename "jit" parameter to "use_jit" for (hopefully) making it self-documenting.

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Remove pytest.mark.parametrize which seems to fail in some circumstances

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Fix unused imports.

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Fix style.

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Give default parameters values for traced model.

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Review comment: Change sentences to sequences

Signed-off-by: Morgan Funtowicz <funtowiczmo@gmail.com>

* Apply suggestions from code review

Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>
2020-11-30 13:43:17 -05:00
LysandreJik 9995a341c9 Update docs 2020-11-30 12:07:52 -05:00
LysandreJik 22b0ff757a Release: v4.0.0 2020-11-30 12:07:43 -05:00
Sylvain GuggerandLysandre Debut 5530299096 Remove deprecated evalutate_during_training (#8852)
* Remove deprecated `evalutate_during_training`

* Update src/transformers/training_args_tf.py

Co-authored-by: Lysandre Debut <lysandre@huggingface.co>

Co-authored-by: Lysandre Debut <lysandre@huggingface.co>
2020-11-30 11:12:15 -05:00
Shai Erera 773849415a Use model.from_pretrained for DataParallel also (#8795)
* Use model.from_pretrained for DataParallel also

When training on multiple GPUs, the code wraps a model with torch.nn.DataParallel. However if the model has custom from_pretrained logic, it does not get applied during load_best_model_at_end.

This commit uses the underlying model during load_best_model_at_end, and re-wraps the loaded model with DataParallel.

If you choose to reject this change, then could you please move the this logic to a function, e.g. def load_best_model_checkpoint(best_model_checkpoint) or something, so that it can be overridden?

* Fix silly bug

* Address review comments

Thanks for the feedback. I made the change that you proposed, but I also think we should update L811 to check if `self.mode` is an instance of `PreTrained`, otherwise we would still not get into that `if` section, right?
2020-11-30 11:11:10 -05:00
Sylvain Gugger 4062c75e44 Merge remote-tracking branch 'origin/master' 2020-11-30 10:51:35 -05:00
Sylvain Gugger 08e707633c Comment the skip job on doc line 2020-11-30 10:51:25 -05:00
Sylvain Gugger 75f8100fc7 Add a direct link to the big table (#8850) 2020-11-30 10:29:23 -05:00
Fraser Greenlee cc983cd9cd Correct docstring. (#8845)
Related issue: https://github.com/huggingface/transformers/issues/8837
2020-11-30 09:33:30 -05:00
Stefan Schweter 19fa01ce2a token-classification: use is_world_process_zero instead of deprecated is_world_master() (#8828) 2020-11-30 09:21:56 -05:00
28 changed files with 277 additions and 96 deletions
+8 -8
View File
@@ -88,7 +88,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-torch_and_tf-{{ checksum "setup.py" }}
@@ -115,7 +115,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-torch-{{ checksum "setup.py" }}
@@ -142,7 +142,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-tf-{{ checksum "setup.py" }}
@@ -169,7 +169,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-flax-{{ checksum "setup.py" }}
@@ -196,7 +196,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-torch-{{ checksum "setup.py" }}
@@ -223,7 +223,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-tf-{{ checksum "setup.py" }}
@@ -248,7 +248,7 @@ jobs:
RUN_CUSTOM_TOKENIZERS: yes
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-custom_tokenizers-{{ checksum "setup.py" }}
@@ -276,7 +276,7 @@ jobs:
parallelism: 1
steps:
- checkout
- skip-job-on-doc-only-changes
# - skip-job-on-doc-only-changes
- restore_cache:
keys:
- v0.4-torch_examples-{{ checksum "setup.py" }}
+2 -1
View File
@@ -52,4 +52,5 @@ deploy_doc "4b3ee9c" v3.1.0
deploy_doc "3ebb1b3" v3.2.0
deploy_doc "0613f05" v3.3.1
deploy_doc "eb0e0ce" v3.4.0
deploy_doc "818878d" # v3.5.1 Latest stable release
deploy_doc "818878d" v3.5.1
deploy_doc "c781171" # v4.0.0 Latest stable release
+2
View File
@@ -198,6 +198,8 @@ ultilingual BERT into [DistilmBERT](https://github.com/huggingface/transformers/
1. **[Other community models](https://huggingface.co/models)**, contributed by the [community](https://huggingface.co/users).
1. Want to contribute a new model? We have added a **detailed guide and templates** to guide you in the process of adding a new model. You can find them in the [`templates`](./templates) folder of the repository. Be sure to check the [contributing guidelines](./CONTRIBUTING.md) and contact the maintainers or open an issue to collect feedbacks before starting your PR.
To cehck if each model has an implementation in PyTorch/TensorFlow/Flax or has an associated tokenizer backed by the 🤗 Tokenizers library, refer to [this table](https://huggingface.co/transformers/index.html#bigtable)
These implementations have been tested on several datasets (see the example scripts) and should match the performances of the original implementations. You can find more details on the performances in the Examples section of the [documentation](https://huggingface.co/transformers/examples.html).
+4 -3
View File
@@ -1,14 +1,15 @@
// These two things need to be updated at each release for the version selector.
// Last stable version
const stableVersion = "v3.5.0"
const stableVersion = "v4.0.0"
// Dictionary doc folder to label
const versionMapping = {
"master": "master",
"": "v3.5.0/v3.5.1",
"v4.0.0": "v4.0.0",
"v3.5.1": "v3.5.0/v3.5.1",
"v3.4.0": "v3.4.0",
"v3.3.1": "v3.3.0/v3.3.1",
"v3.2.0": "v3.2.0",
"v3.1.0": "v3.1.0 (stable)",
"v3.1.0": "v3.1.0",
"v3.0.2": "v3.0.0/v3.0.1/v3.0.2",
"v2.11.0": "v2.11.0",
"v2.10.0": "v2.10.0",
+1 -1
View File
@@ -26,7 +26,7 @@ author = u'huggingface'
# The short X.Y version
version = u''
# The full version, including alpha/beta/rc tags
release = u'3.5.0'
release = u'4.0.0'
# -- General configuration ---------------------------------------------------
+2
View File
@@ -172,6 +172,8 @@ and conversion utilities for the following models:
<https://huggingface.co/users>`__.
.. _bigtable:
The table below represents the current support in the library for each of those models, whether they have a Python
tokenizer (called "slow"). A "fast" tokenizer backed by the 🤗 Tokenizers library, whether they have support in PyTorch,
TensorFlow and/or Flax.
+2 -1
View File
@@ -3,7 +3,8 @@
python finetune_trainer.py \
--learning_rate=3e-5 \
--fp16 \
--do_train --do_eval --do_predict --evaluate_during_training \
--do_train --do_eval --do_predict \
--evaluation_strategy steps \
--predict_with_generate \
--n_val 1000 \
"$@"
@@ -5,7 +5,8 @@ export TPU_NUM_CORES=8
python xla_spawn.py --num_cores $TPU_NUM_CORES \
finetune_trainer.py \
--learning_rate=3e-5 \
--do_train --do_eval --evaluate_during_training \
--do_train --do_eval \
--evaluation_strategy steps \
--prediction_loss_only \
--n_val 1000 \
"$@"
@@ -16,7 +16,8 @@ python finetune_trainer.py \
--num_train_epochs=6 \
--save_steps 3000 --eval_steps 3000 \
--max_source_length $MAX_LEN --max_target_length $MAX_LEN --val_max_target_length $MAX_LEN --test_max_target_length $MAX_LEN \
--do_train --do_eval --do_predict --evaluate_during_training\
--do_train --do_eval --do_predict \
--evaluation_strategy steps \
--predict_with_generate --logging_first_step \
--task translation --label_smoothing 0.1 \
"$@"
@@ -17,7 +17,8 @@ python xla_spawn.py --num_cores $TPU_NUM_CORES \
--save_steps 500 --eval_steps 500 \
--logging_first_step --logging_steps 200 \
--max_source_length $MAX_LEN --max_target_length $MAX_LEN --val_max_target_length $MAX_LEN --test_max_target_length $MAX_LEN \
--do_train --do_eval --evaluate_during_training \
--do_train --do_eval \
--evaluation_strategy steps \
--prediction_loss_only \
--task translation --label_smoothing 0.1 \
"$@"
@@ -19,6 +19,7 @@ python finetune_trainer.py \
--save_steps 3000 --eval_steps 3000 \
--logging_first_step \
--max_target_length 56 --val_max_target_length $MAX_TGT_LEN --test_max_target_length $MAX_TGT_LEN \
--do_train --do_eval --do_predict --evaluate_during_training \
--do_train --do_eval --do_predict \
--evaluation_strategy steps \
--predict_with_generate --sortish_sampler \
"$@"
@@ -15,7 +15,8 @@ python finetune_trainer.py \
--sortish_sampler \
--num_train_epochs 6 \
--save_steps 25000 --eval_steps 25000 --logging_steps 1000 \
--do_train --do_eval --do_predict --evaluate_during_training \
--do_train --do_eval --do_predict \
--evaluation_strategy steps \
--predict_with_generate --logging_first_step \
--task translation \
"$@"
+2 -2
View File
@@ -369,7 +369,7 @@ def main():
]
output_test_results_file = os.path.join(training_args.output_dir, "test_results.txt")
if trainer.is_world_master():
if trainer.is_world_process_zero():
with open(output_test_results_file, "w") as writer:
for key, value in metrics.items():
logger.info(f" {key} = {value}")
@@ -377,7 +377,7 @@ def main():
# Save predictions
output_test_predictions_file = os.path.join(training_args.output_dir, "test_predictions.txt")
if trainer.is_world_master():
if trainer.is_world_process_zero():
with open(output_test_predictions_file, "w") as writer:
for prediction in true_predictions:
writer.write(" ".join(prediction) + "\n")
+2 -2
View File
@@ -291,7 +291,7 @@ def main():
preds_list, _ = align_predictions(predictions, label_ids)
output_test_results_file = os.path.join(training_args.output_dir, "test_results.txt")
if trainer.is_world_master():
if trainer.is_world_process_zero():
with open(output_test_results_file, "w") as writer:
for key, value in metrics.items():
logger.info(" %s = %s", key, value)
@@ -299,7 +299,7 @@ def main():
# Save predictions
output_test_predictions_file = os.path.join(training_args.output_dir, "test_predictions.txt")
if trainer.is_world_master():
if trainer.is_world_process_zero():
with open(output_test_predictions_file, "w") as writer:
with open(os.path.join(data_args.data_dir, "test.txt"), "r") as f:
token_classification_task.write_predictions_to_file(writer, f, preds_list)
+1 -1
View File
@@ -230,7 +230,7 @@ install_requires = [
setup(
name="transformers",
version="4.0.0-rc-1",
version="4.1.0.dev0",
author="Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Sam Shleifer, Patrick von Platen, Sylvain Gugger, Google AI Language Team Authors, Open AI team Authors, Facebook AI Authors, Carnegie Mellon University Authors",
author_email="thomas@huggingface.co",
description="State-of-the-art Natural Language Processing for TensorFlow 2.0 and PyTorch",
+1 -1
View File
@@ -2,7 +2,7 @@
# There's no way to ignore "F401 '...' imported but unused" warnings in this
# module, but to preserve other warnings. So, don't check this module at all.
__version__ = "4.0.0-rc-1"
__version__ = "4.1.0.dev0"
# Work around to update TensorFlow's absl.logging threshold which alters the
# default Python logging output behavior when present.
+3 -2
View File
@@ -2,6 +2,7 @@
import math
import os
from .trainer_utils import EvaluationStrategy
from .utils import logging
@@ -212,13 +213,13 @@ def run_hp_search_ray(trainer, n_trials: int, direction: str, **kwargs) -> BestR
# Check for `do_eval` and `eval_during_training` for schedulers that require intermediate reporting.
if isinstance(
kwargs["scheduler"], (ASHAScheduler, MedianStoppingRule, HyperBandForBOHB, PopulationBasedTraining)
) and (not trainer.args.do_eval or not trainer.args.evaluate_during_training):
) and (not trainer.args.do_eval or trainer.args.evaluation_strategy == EvaluationStrategy.NO):
raise RuntimeError(
"You are using {cls} as a scheduler but you haven't enabled evaluation during training. "
"This means your trials will not report intermediate results to Ray Tune, and "
"can thus not be stopped early or used to exploit other trials parameters. "
"If this is what you want, do not use {cls}. If you would like to use {cls}, "
"make sure you pass `do_eval=True` and `evaluate_during_training=True` in the "
"make sure you pass `do_eval=True` and `evaluation_strategy='steps'` in the "
"Trainer `args`.".format(cls=type(kwargs["scheduler"]).__name__)
)
@@ -459,6 +459,7 @@ class LongformerEmbeddings(nn.Module):
# position_ids (1, len position emb) is contiguous in memory and exported when serialized
self.register_buffer("position_ids", torch.arange(config.max_position_embeddings).expand((1, -1)))
self.position_embedding_type = getattr(config, "position_embedding_type", "absolute")
self.padding_idx = config.pad_token_id
self.position_embeddings = nn.Embedding(
+71 -10
View File
@@ -516,7 +516,7 @@ class Pipeline(_ScikitCompat):
if framework is None:
framework = get_framework(model)
self.task = task
self.task = task or getattr(self, "task", "")
self.model = model
self.tokenizer = tokenizer
self.modelcard = modelcard
@@ -530,8 +530,8 @@ class Pipeline(_ScikitCompat):
# Update config with task specific parameters
task_specific_params = self.model.config.task_specific_params
if task_specific_params is not None and task in task_specific_params:
self.model.config.update(task_specific_params.get(task))
if task_specific_params is not None and self.task in task_specific_params:
self.model.config.update(task_specific_params.get(self.task))
def save_pretrained(self, save_directory: str):
"""
@@ -926,6 +926,14 @@ class TextGenerationPipeline(Pipeline):
r"""
return_all_scores (:obj:`bool`, `optional`, defaults to :obj:`False`):
Whether to return all prediction scores or just the one of the predicted class.
function_to_apply (:obj:`str`, `optional`, defaults to :obj:`"default"`):
The function to apply to the model outputs in order to retrieve the scores. Accepts four different values:
- :obj:`"default"`: if the model has a single label, will apply the sigmoid function on the output. If the
model has several labels, will apply the softmax function on the output.
- :obj:`"sigmoid"`: Applies the sigmoid function on the output.
- :obj:`"softmax"`: Applies the softmax function on the output.
- :obj:`"none"`: Does not apply any function on the output.
""",
)
class TextClassificationPipeline(Pipeline):
@@ -945,7 +953,9 @@ class TextClassificationPipeline(Pipeline):
<https://huggingface.co/models?filter=text-classification>`__.
"""
def __init__(self, return_all_scores: bool = False, **kwargs):
task = "text-classification"
def __init__(self, return_all_scores: bool = None, function_to_apply: str = None, **kwargs):
super().__init__(**kwargs)
self.check_model_type(
@@ -954,15 +964,33 @@ class TextClassificationPipeline(Pipeline):
else MODEL_FOR_SEQUENCE_CLASSIFICATION_MAPPING
)
self.return_all_scores = return_all_scores
if hasattr(self.model.config, "return_all_scores") and return_all_scores is None:
return_all_scores = self.model.config.return_all_scores
def __call__(self, *args, **kwargs):
if hasattr(self.model.config, "function_to_apply") and function_to_apply is None:
function_to_apply = self.model.config.function_to_apply
self.return_all_scores = return_all_scores if return_all_scores is not None else False
self.function_to_apply = function_to_apply if function_to_apply is not None else "default"
def __call__(self, *args, return_all_scores=None, function_to_apply=None, **kwargs):
"""
Classify the text(s) given as inputs.
Args:
args (:obj:`str` or :obj:`List[str]`):
One or several texts (or one list of prompts) to classify.
return_all_scores (:obj:`bool`, `optional`, defaults to :obj:`False`):
Whether to return scores for all labels.
function_to_apply (:obj:`str`, `optional`, defaults to :obj:`"default"`):
The function to apply to the model outputs in order to retrieve the scores. Accepts four different
values:
- :obj:`"default"`: if the model has a single label, will apply the sigmoid function on the output. If
the model has several labels, will apply the softmax function on the output.
- :obj:`"sigmoid"`: Applies the sigmoid function on the output.
- :obj:`"softmax"`: Applies the softmax function on the output.
- :obj:`"none"`: Does not apply any function on the output.
Return:
A list or a list of list of :obj:`dict`: Each result comes as list of dictionaries with the following keys:
@@ -974,11 +1002,30 @@ class TextClassificationPipeline(Pipeline):
"""
outputs = super().__call__(*args, **kwargs)
if self.model.config.num_labels == 1:
scores = 1.0 / (1.0 + np.exp(-outputs))
return_all_scores = return_all_scores if return_all_scores is not None else self.return_all_scores
function_to_apply = function_to_apply if function_to_apply is not None else self.function_to_apply
def sigmoid(_outputs):
return 1.0 / (1.0 + np.exp(-_outputs))
def softmax(_outputs):
return np.exp(_outputs) / np.exp(_outputs).sum(-1, keepdims=True)
if function_to_apply == "default":
if self.model.config.num_labels == 1:
scores = sigmoid(outputs)
else:
scores = softmax(outputs)
elif function_to_apply == "sigmoid":
scores = sigmoid(outputs)
elif function_to_apply == "softmax":
scores = softmax(outputs)
elif function_to_apply.lower() == "none":
scores = outputs
else:
scores = np.exp(outputs) / np.exp(outputs).sum(-1, keepdims=True)
if self.return_all_scores:
raise ValueError(f"Unrecognized `function_to_apply` argument: {function_to_apply}")
if return_all_scores:
return [
[{"label": self.model.config.id2label[i], "score": score.item()} for i, score in enumerate(item)]
for item in scores
@@ -1040,6 +1087,8 @@ class ZeroShotClassificationPipeline(Pipeline):
of available models on `huggingface.co/models <https://huggingface.co/models?search=nli>`__.
"""
task = "zero-shot"
def __init__(self, args_parser=ZeroShotClassificationArgumentHandler(), *args, **kwargs):
super().__init__(*args, **kwargs)
self._args_parser = args_parser
@@ -1172,6 +1221,8 @@ class FillMaskPipeline(Pipeline):
This pipeline only works for inputs with exactly one token masked.
"""
task = "fill-mask"
def __init__(
self,
model: Union["PreTrainedModel", "TFPreTrainedModel"],
@@ -1361,6 +1412,7 @@ class TokenClassificationPipeline(Pipeline):
<https://huggingface.co/models?filter=token-classification>`__.
"""
task = "token-classification"
default_input_names = "sequences"
def __init__(
@@ -1667,6 +1719,7 @@ class QuestionAnsweringPipeline(Pipeline):
<https://huggingface.co/models?filter=question-answering>`__.
"""
task = "question-answering"
default_input_names = "question,context"
def __init__(
@@ -2067,6 +2120,8 @@ class SummarizationPipeline(Pipeline):
summarizer("Sam Shleifer writes the best docstring examples in the whole world.", min_length=5, max_length=20)
"""
task = "summarization"
def __init__(self, *args, **kwargs):
kwargs.update(task="summarization")
super().__init__(*args, **kwargs)
@@ -2188,6 +2243,8 @@ class TranslationPipeline(Pipeline):
en_fr_translator("How old are you?")
"""
task = "translation"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
@@ -2297,6 +2354,8 @@ class Text2TextGenerationPipeline(Pipeline):
text2text_generator("question: What is 42 ? context: 42 is the answer to life, the universe and everything")
"""
task = "text2text-generation"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
@@ -2522,6 +2581,8 @@ class ConversationalPipeline(Pipeline):
conversational_pipeline([conversation_1, conversation_2])
"""
task = "conversational"
def __init__(self, min_length_for_response=32, *args, **kwargs):
super().__init__(*args, **kwargs)
+1 -1
View File
@@ -2604,7 +2604,7 @@ class PreTrainedTokenizerBase(SpecialTokensMixin):
Instead of :obj:`List[int]` you can have tensors (numpy arrays, PyTorch tensors or TensorFlow tensors),
see the note above for the return type.
padding (:obj:`bool`, :obj:`str` or :class:`~transformers.tokenization_utils_base.PaddingStrategy`, `optional`, defaults to :obj:`False`):
padding (:obj:`bool`, :obj:`str` or :class:`~transformers.tokenization_utils_base.PaddingStrategy`, `optional`, defaults to :obj:`True`):
Select a strategy to pad the returned sequences (according to the model's padding side and padding
index) among:
+2 -2
View File
@@ -808,8 +808,8 @@ class Trainer:
logger.info(
f"Loading best model from {self.state.best_model_checkpoint} (score: {self.state.best_metric})."
)
if isinstance(model, PreTrainedModel):
self.model = model.from_pretrained(self.state.best_model_checkpoint)
if isinstance(self.model, PreTrainedModel):
self.model = self.model.from_pretrained(self.state.best_model_checkpoint)
if not self.args.model_parallel:
self.model = self.model.to(self.args.device)
else:
+2 -2
View File
@@ -19,7 +19,7 @@ from tensorflow.python.distribute.values import PerReplica
from .modeling_tf_utils import TFPreTrainedModel
from .optimization_tf import GradientAccumulator, create_optimizer
from .trainer_utils import PREFIX_CHECKPOINT_DIR, EvalPrediction, PredictionOutput, set_seed
from .trainer_utils import PREFIX_CHECKPOINT_DIR, EvalPrediction, EvaluationStrategy, PredictionOutput, set_seed
from .training_args_tf import TFTrainingArguments
from .utils import logging
@@ -561,7 +561,7 @@ class TFTrainer:
if (
self.args.eval_steps > 0
and self.args.evaluate_during_training
and self.args.evaluate_strategy == EvaluationStrategy.STEPS
and self.global_step % self.args.eval_steps == 0
):
self.evaluate()
+6 -2
View File
@@ -34,8 +34,12 @@ class TFTrainingArguments(TrainingArguments):
Whether to run evaluation on the dev set or not.
do_predict (:obj:`bool`, `optional`, defaults to :obj:`False`):
Whether to run predictions on the test set or not.
evaluate_during_training (:obj:`bool`, `optional`, defaults to :obj:`False`):
Whether to run evaluation during training at each logging step or not.
evaluation_strategy (:obj:`str` or :class:`~transformers.trainer_utils.EvaluationStrategy`, `optional`, defaults to :obj:`"no"`):
The evaluation strategy to adopt during training. Possible values are:
* :obj:`"no"`: No evaluation is done during training.
* :obj:`"steps"`: Evaluation is done (and logged) every :obj:`eval_steps`.
per_device_train_batch_size (:obj:`int`, `optional`, defaults to 8):
The batch size per GPU/TPU core/CPU for training.
per_device_eval_batch_size (:obj:`int`, `optional`, defaults to 8):
+22 -24
View File
@@ -1,6 +1,5 @@
import unittest
import pytest
from numpy import ndarray
from transformers import BertTokenizerFast, TensorType, is_flax_available, is_torch_available
@@ -24,6 +23,10 @@ if is_torch_available():
@require_flax
@require_torch
class FlaxBertModelTest(unittest.TestCase):
def assert_almost_equals(self, a: ndarray, b: ndarray, tol: float):
diff = (a - b).sum()
self.assertLessEqual(diff, tol, f"Difference between torch and flax is {diff} (>= {tol})")
def test_from_pytorch(self):
with torch.no_grad():
with self.subTest("bert-base-cased"):
@@ -40,32 +43,27 @@ class FlaxBertModelTest(unittest.TestCase):
self.assertEqual(len(fx_outputs), len(pt_outputs), "Output lengths differ between Flax and PyTorch")
for fx_output, pt_output in zip(fx_outputs, pt_outputs):
self.assert_almost_equals(fx_output, pt_output.numpy(), 5e-4)
self.assert_almost_equals(fx_output, pt_output.numpy(), 5e-3)
def assert_almost_equals(self, a: ndarray, b: ndarray, tol: float):
diff = (a - b).sum()
self.assertLessEqual(diff, tol, "Difference between torch and flax is {} (>= {})".format(diff, tol))
def test_multiple_sequences(self):
tokenizer = BertTokenizerFast.from_pretrained("bert-base-cased")
model = FlaxBertModel.from_pretrained("bert-base-cased")
sequences = ["this is an example sentence", "this is another", "and a third one"]
encodings = tokenizer(sequences, return_tensors=TensorType.JAX, padding=True, truncation=True)
@require_flax
@require_torch
@pytest.mark.parametrize("jit", ["disable_jit", "enable_jit"])
def test_multiple_sentences(jit):
tokenizer = BertTokenizerFast.from_pretrained("bert-base-cased")
model = FlaxBertModel.from_pretrained("bert-base-cased")
@jax.jit
def model_jitted(input_ids, attention_mask=None, token_type_ids=None):
return model(input_ids, attention_mask, token_type_ids)
sentences = ["this is an example sentence", "this is another", "and a third one"]
encodings = tokenizer(sentences, return_tensors=TensorType.JAX, padding=True, truncation=True)
with self.subTest("JIT Disabled"):
with jax.disable_jit():
tokens, pooled = model_jitted(**encodings)
self.assertEqual(tokens.shape, (3, 7, 768))
self.assertEqual(pooled.shape, (3, 768))
@jax.jit
def model_jitted(input_ids, attention_mask, token_type_ids):
return model(input_ids, attention_mask, token_type_ids)
with self.subTest("JIT Enabled"):
jitted_tokens, jitted_pooled = model_jitted(**encodings)
if jit == "disable_jit":
with jax.disable_jit():
tokens, pooled = model_jitted(**encodings)
else:
tokens, pooled = model_jitted(**encodings)
assert tokens.shape == (3, 7, 768)
assert pooled.shape == (3, 768)
self.assertEqual(jitted_tokens.shape, (3, 7, 768))
self.assertEqual(jitted_pooled.shape, (3, 768))
+22 -24
View File
@@ -1,6 +1,5 @@
import unittest
import pytest
from numpy import ndarray
from transformers import RobertaTokenizerFast, TensorType, is_flax_available, is_torch_available
@@ -24,6 +23,10 @@ if is_torch_available():
@require_flax
@require_torch
class FlaxRobertaModelTest(unittest.TestCase):
def assert_almost_equals(self, a: ndarray, b: ndarray, tol: float):
diff = (a - b).sum()
self.assertLessEqual(diff, tol, f"Difference between torch and flax is {diff} (>= {tol})")
def test_from_pytorch(self):
with torch.no_grad():
with self.subTest("roberta-base"):
@@ -40,32 +43,27 @@ class FlaxRobertaModelTest(unittest.TestCase):
self.assertEqual(len(fx_outputs), len(pt_outputs), "Output lengths differ between Flax and PyTorch")
for fx_output, pt_output in zip(fx_outputs, pt_outputs.to_tuple()):
self.assert_almost_equals(fx_output, pt_output.numpy(), 5e-4)
self.assert_almost_equals(fx_output, pt_output.numpy(), 6e-4)
def assert_almost_equals(self, a: ndarray, b: ndarray, tol: float):
diff = (a - b).sum()
self.assertLessEqual(diff, tol, "Difference between torch and flax is {} (>= {})".format(diff, tol))
def test_multiple_sequences(self):
tokenizer = RobertaTokenizerFast.from_pretrained("roberta-base")
model = FlaxRobertaModel.from_pretrained("roberta-base")
sequences = ["this is an example sentence", "this is another", "and a third one"]
encodings = tokenizer(sequences, return_tensors=TensorType.JAX, padding=True, truncation=True)
@require_flax
@require_torch
@pytest.mark.parametrize("jit", ["disable_jit", "enable_jit"])
def test_multiple_sentences(jit):
tokenizer = RobertaTokenizerFast.from_pretrained("roberta-base")
model = FlaxRobertaModel.from_pretrained("roberta-base")
@jax.jit
def model_jitted(input_ids, attention_mask=None, token_type_ids=None):
return model(input_ids, attention_mask, token_type_ids)
sentences = ["this is an example sentence", "this is another", "and a third one"]
encodings = tokenizer(sentences, return_tensors=TensorType.JAX, padding=True, truncation=True)
with self.subTest("JIT Disabled"):
with jax.disable_jit():
tokens, pooled = model_jitted(**encodings)
self.assertEqual(tokens.shape, (3, 7, 768))
self.assertEqual(pooled.shape, (3, 768))
@jax.jit
def model_jitted(input_ids, attention_mask):
return model(input_ids, attention_mask)
with self.subTest("JIT Enabled"):
jitted_tokens, jitted_pooled = model_jitted(**encodings)
if jit == "disable_jit":
with jax.disable_jit():
tokens, pooled = model_jitted(**encodings)
else:
tokens, pooled = model_jitted(**encodings)
assert tokens.shape == (3, 7, 768)
assert pooled.shape == (3, 768)
self.assertEqual(jitted_tokens.shape, (3, 7, 768))
self.assertEqual(jitted_pooled.shape, (3, 768))
+1
View File
@@ -197,6 +197,7 @@ class MonoInputPipelineCommonMixin(CustomInputPipelineCommonMixin):
def _test_pipeline(self, nlp: Pipeline):
self.assertIsNotNone(nlp)
self.assertIsNotNone(nlp.task)
mono_result = nlp(self.valid_inputs[0], **self.pipeline_running_kwargs)
self.assertIsInstance(mono_result, list)
+108 -3
View File
@@ -1,12 +1,117 @@
import unittest
import numpy as np
from transformers import AutoTokenizer, DistilBertConfig, DistilBertForSequenceClassification, pipeline
from transformers.testing_utils import slow
from .test_pipelines_common import MonoInputPipelineCommonMixin
VALID_INPUTS = ["I really disagree with what you've said.", ["I love you."]]
class SentimentAnalysisPipelineTests(MonoInputPipelineCommonMixin, unittest.TestCase):
pipeline_task = "sentiment-analysis"
small_models = [
"sshleifer/tiny-distilbert-base-uncased-finetuned-sst-2-english"
] # Default model - Models tested without the @slow decorator
small_models = ["distilbert-base-cased"] # Default model - Models tested without the @slow decorator
large_models = [None] # Models tested with the @slow decorator
mandatory_keys = {"label", "score"} # Keys which should be in the output
@slow
def test_function_to_apply(self):
for model_name in self.small_models:
string_input, string_list_input = VALID_INPUTS
tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=True)
model = DistilBertForSequenceClassification(DistilBertConfig.from_pretrained(model_name))
model.eval()
classifier = pipeline(task="sentiment-analysis", model=model, tokenizer=tokenizer)
def check_output(output):
# This model does not have a sequence-classification head, so results are random
self.assertTrue(isinstance(output, list))
self.assertEqual(len(output), 1)
d = string_output[0]
self.assertEqual(set(d.keys()), {"label", "score"})
self.assertEqual(type(d["label"]), str)
self.assertEqual(type(d["score"]), float)
def pipeline_call_argument(argument=None):
_string_output = classifier(string_input, function_to_apply=argument)
_string_list_output = classifier(string_list_input, function_to_apply=argument)
check_output(_string_output)
return _string_output, _string_list_output
def pipeline_init_argument(argument=None):
_classifier = pipeline(
task="sentiment-analysis", model=model, tokenizer=tokenizer, function_to_apply=argument
)
_string_output = _classifier(string_input)
_string_list_output = _classifier(string_list_input)
check_output(_string_output)
return _string_output, _string_list_output
def pipeline_model_argument(argument=None):
model.config.task_specific_params = {"sentiment-analysis": {"function_to_apply": argument}}
_classifier = pipeline(task="sentiment-analysis", model=model, tokenizer=tokenizer)
_string_output = _classifier(string_input)
_string_list_output = _classifier(string_list_input)
check_output(_string_output)
return _string_output, _string_list_output
string_output = classifier(string_input)
string_list_output = classifier(string_list_input)
string_output_default = pipeline_call_argument("default")
string_output_sigmoid = pipeline_call_argument("sigmoid")
string_output_softmax = pipeline_call_argument("softmax")
string_output_none = pipeline_call_argument("none")
string_output_init_default = pipeline_init_argument("default")
string_output_init_sigmoid = pipeline_init_argument("sigmoid")
string_output_init_softmax = pipeline_init_argument("softmax")
string_output_init_none = pipeline_init_argument("none")
string_output_model_default = pipeline_model_argument("default")
string_output_model_sigmoid = pipeline_model_argument("sigmoid")
string_output_model_softmax = pipeline_model_argument("softmax")
string_output_model_none = pipeline_model_argument("none")
should_be_equal = [
(
(string_output, string_list_output),
string_output_default,
string_output_init_default,
string_output_model_default,
),
(string_output_sigmoid, string_output_init_sigmoid, string_output_model_sigmoid),
(string_output_softmax, string_output_init_softmax, string_output_model_softmax),
(string_output_none, string_output_init_none, string_output_model_none),
]
for tuples_containing_equal_values in should_be_equal:
# Retrieve each tuple from the list
for tuple_value_0 in tuples_containing_equal_values:
for tuple_value_1 in tuples_containing_equal_values:
if tuple_value_0 is not tuple_value_1:
# Compare each tuple value with all the others, as long as they're not the same object
for pipeline_output_0, pipeline_output_1 in zip(tuple_value_0, tuple_value_1):
# Compare all outputs (call, init, model argument)
for example_result_0, example_result_1 in zip(pipeline_output_0, pipeline_output_1):
# Iterate through the results
self.assertTrue(
np.allclose(example_result_0["score"], example_result_1["score"], atol=1e-6)
)
@slow
def test_function_to_apply_error(self):
for model_name in self.small_models:
string_input, _ = VALID_INPUTS
tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=True)
model = DistilBertForSequenceClassification(DistilBertConfig.from_pretrained(model_name))
model.eval()
classifier = pipeline(task="sentiment-analysis", model=model, tokenizer=tokenizer)
with self.assertRaises(ValueError):
classifier(string_input, function_to_apply="logits")
+1 -1
View File
@@ -282,7 +282,7 @@ def check_model_list_copy(overwrite=False, max_per_line=119):
rst_list, start_index, end_index, lines = _find_text_in_file(
filename=os.path.join(PATH_TO_DOCS, "index.rst"),
start_prompt=" This list is updated automatically from the README",
end_prompt="The table below represents the current support",
end_prompt=".. _bigtable:",
)
md_list = get_model_list()
converted_list = convert_to_rst(md_list, max_per_line=max_per_line)