Joe Davison and Patrick von Platen
369f1d77b4
Return correct Bart hidden state tensors ( #8747 )
...
* bart output hidden states upstream
* same w/ decoder
* add tests
* fix prophetnet
* fix gpt2 and ctrl
* fix fstm and skip test for reformer and longformer
* fix all models
Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com >
2020-11-25 22:06:04 +01:00
Joe Davison
f6f4da8dd4
Add bart-large-mnli model card ( #8527 )
2020-11-13 14:07:25 -05:00
Joe Davison
556709ad92
rm multiclass option from model card
2020-10-27 17:11:43 -04:00
Joe Davison
3e58b6b7b8
infer entailment label id on zero shot pipeline ( #8059 )
...
* add entailment dim argument
* rename dim -> id
* fix last name change, style
* rm arg, auto-infer only
* typo
* rm superfluous import
2020-10-27 14:09:55 -04:00
Joe Davison
fbcddb8544
add mutliclass field to default zero shot example
2020-10-26 11:07:51 -04:00
Joe Davison
b0a907615a
minor model card description updates ( #8051 )
2020-10-26 10:04:20 -04:00
Joe Davison
64b24bb3c2
change zero shot widget default example ( #7992 )
2020-10-22 15:19:41 -06:00
Joe Davison and Julien Chaumond
077c99bb5f
add zero shot pipeline tags & examples ( #7983 )
...
* add zero shot pipeline tags
* rm default and fix yaml format
* rm DS_Store
* add bart large default
* don't add more typos
Co-authored-by: Julien Chaumond <chaumond@gmail.com >
* add multiple multilingual examples
* improve multilingual examples for single-label
Co-authored-by: Julien Chaumond <chaumond@gmail.com >
2020-10-22 13:01:23 -06:00
Joe Davison
13842e413c
PPL guide minor code snippet fix ( #7938 )
2020-10-20 16:17:39 -06:00
Joe Davison
a1ac082879
add license to xlm-roberta-large-xnli card
2020-10-09 09:16:06 -04:00
Joe Davison
10a34501f1
add __init__.py to utils ( #6754 )
2020-08-26 23:51:10 +02:00
Joe Davison
99407f9d1e
add xlm-roberta-large-xnli model card ( #6723 )
...
* add xlm-roberta-large-xnli model card
* update pt example
* typo
2020-08-26 16:05:59 -04:00
Joe Davison
f9d280a959
TFTrainer dataset doc & fix evaluation bug ( #6618 )
...
* TFTrainer dataset doc & fix evaluation bug
discussed in #6551
* add docstring to test/eval datasets
2020-08-20 12:11:36 -04:00
Joe Davison
039d8d65fc
add intro to nlp lib & dataset links to custom datasets tutorial ( #6583 )
...
* add intro to nlp lib + links
* unique links...
2020-08-20 10:32:51 -04:00
Joe Davison and Sylvain Gugger
d0c2389f48
add custom datasets tutorial ( #6466 )
...
* add custom datasets tutorial
* python -> bash code blocks
* Apply suggestions from code review
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
* minor review feedback changes
* add working native QA snippet
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
2020-08-17 09:15:34 -04:00
Joe Davison
bc820476a5
add targets arg to fill-mask pipeline ( #6239 )
...
* add targets arg to fill-mask pipeline
* add tests and more error handling
* quality
* update docstring
2020-08-12 12:48:29 -04:00
Joe Davison
972535ea74
fix zero shot pipeline docs ( #6245 )
2020-08-04 16:37:49 -04:00
Joe Davison and Julien Chaumond
8edfaaa81b
bart-large-mnli-yahoo-answers model card ( #6133 )
...
* Add bart-large-mnli-yahoo-answers model card
* Add examples
* Add widget example
* Rm bart tag
Co-authored-by: Julien Chaumond <chaumond@gmail.com >
Co-authored-by: Julien Chaumond <chaumond@gmail.com >
2020-07-31 10:56:32 -04:00
Joe Davison
b1c8b76907
Fix zero-shot pipeline single seq output shape ( #6104 )
2020-07-28 14:46:03 -04:00
Joe Davison
3deffc1d67
Zero shot classification pipeline ( #5760 )
...
* add initial zero-shot pipeline
* change default args
* update default template
* add label string splitting
* add str labels support, remove nli from name
* style
* add input validation and working tf defaults
* tests
* quality check
* add docstring to __call__
* add slow tests
* Change truncation to only_first
also lower precision on tests for readibility
* style
2020-07-27 09:42:58 -04:00
Joe Davison
5d178954c9
tiny ppl doc typo fix ( #5751 )
2020-07-14 10:39:44 -06:00
Joe Davison
b4b33fdf25
Guide to fixed-length model perplexity evaluation ( #5449 )
...
* add first draft ppl guide
* upload imgs
* expand on strides
* ref typo
* rm superfluous past var
* add tokenization disclaimer
2020-07-07 16:04:15 -06:00
Joe Davison
35befd9ce3
Fix tensor label type inference in default collator ( #5250 )
...
* allow tensor label inputs to default collator
* replace try/except with type check
2020-07-01 10:40:14 -06:00
Joe Davison and Sylvain Gugger
2ffef0d0c7
Training & fine-tuning quickstart ( #5034 )
...
* add initial fine-tuning guide
* split code blocks to smaller segments
* fix up trianer section of fine-tune doc
* a few last typos
* Update usage -> task summary link
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
2020-06-25 15:11:11 -06:00
Joe Davison
c36416e53c
Add standardized get_vocab method to tokenizers
2020-02-22 12:09:01 -05:00
Joe Davison
197d74f988
Add get_vocab method to PretrainedTokenizer
2020-02-20 15:26:49 -05:00
Joe Davison
f1e8a51f08
Preserve spaces in GPT-2 tokenizers ( #2778 )
...
* Preserve spaces in GPT-2 tokenizers
Preserves spaces after special tokens in GPT-2 and inhereted (RoBERTa)
tokenizers, enabling correct BPE encoding. Automatically inserts a space
in front of first token in encode function when adding special tokens.
* Add tokenization preprocessing method
* Add framework argument to pipeline factory
Also fixes pipeline test issue. Each test input now treated as a
distinct sequence.
2020-02-13 13:29:43 -05:00