Patrick von Platen
6c091abef2
[Templates] Adapt Bert ( #9284 )
...
* adapt templates
* adapt config
* add test as well
* fix output type
* fix cache false naming
* finish tests
* last fix
2020-12-24 01:44:33 +01:00
Patrick von Platen
d5db6c37d4
[Seq2Seq Templates] Fix check_repo.py templates file ( #9277 )
...
* add enc dec pt model to check repo
* fix indent
2020-12-23 11:40:20 +01:00
Patrick von Platen
cbe63949d7
Model Templates for Seq2Seq ( #9251 )
...
* adapt cookie cutter
* fix copy past statement
* delete copy statements for now
* remove unused import from template
* make doc rst
* correct config docstring
* correct training
* correct inputs processing tf enc dec
* make style
* adapt templates
* clean tabs
* correct tensor -> Tensor naming
* correct indent
* correct templates
* fix the test
* break lines to avoid > 119
* Apply suggestions from code review
2020-12-22 23:41:20 +01:00
Patrick von Platen
e9d77ccd5a
[EncoderDecoder] Make tests more aggressive ( #9256 )
...
* add tests
* make style and fix bart bug
* fix bart past key value edge case
* correct tf bart test
* fix gpt2 tf
* fix t5 test
2020-12-22 17:00:04 +01:00
Patrick von Platen
9a12b9696f
[MPNet] Add slow to fast tokenizer converter ( #9233 )
...
* add converter
* delet unnecessary comments
2020-12-21 15:41:34 +01:00
Patrick von Platen
6b034309ca
fix warning ( #9231 )
2020-12-21 10:41:34 +01:00
Patrick von Platen and TevenLeScao
640e6fe190
[Flax] Align FlaxBertForMaskedLM with BertForMaskedLM, implement from_pretrained, init ( #9054 )
...
* save intermediate
* save intermediate
* save intermediate
* correct flax bert model file
* new module / model naming
* make style
* almost finish BERT
* finish roberta
* make fix-copies
* delete keys file
* last refactor
* fixes in run_mlm_flax.py
* remove pooled from run_mlm_flax.py`
* fix gelu | gelu_new
* remove Module from inits
* splits
* dirty print
* preventing warmup_steps == 0
* smaller splits
* make fix-copies
* dirty print
* dirty print
* initial_evaluation argument
* declaration order fix
* proper model initialization/loading
* proper initialization
* run_mlm_flax improvements: improper model inputs bugfix + automatic dataset splitting + tokenizers parallelism warning + avoiding warmup_steps=0 bug
* removed tokenizers warning hack, fixed model re-initialization
* reverted training_args.py changes
* fix flax from pretrained
* improve test in flax
* apply sylvains tips
* update init
* make 0.3.0 compatible
* revert tevens changes
* revert tevens changes 2
* finalize revert
* fix bug
* add docs
* add pretrained to init
* Update src/transformers/modeling_flax_utils.py
* fix copies
* final improvements
Co-authored-by: TevenLeScao <teven.lescao@gmail.com >
2020-12-16 13:03:32 +01:00
Patrick von Platen
18ecd36f65
Fix Bart Shift ( #9135 )
...
* correct mistake in order
* fix tensor copy
* clone tensor correctly
2020-12-15 19:04:31 +01:00
Patrick von Platen
d018622d8e
correct mistake in order ( #9134 )
2020-12-15 23:08:31 +05:30
Patrick von Platen
80bdb9c31a
fix bart loss masking ( #9131 )
2020-12-15 18:17:17 +01:00
Patrick von Platen
abc573f51a
[TF Bart] Refactor TFBart ( #9029 )
...
* reorder file
* delete unnecesarry function
* make style
* save intermediate
* fix attention masks
* correct tf bart past key values
* solve merge conflict bug
* correct tensor dims
* save intermediate tf
* change attn layer
* fix typo re-order past
* inputs_embeds
* make fix copies
* finish tests
* fix graph mode
* appyl lysandres suggestions
2020-12-15 17:31:28 +01:00
Patrick von Platen
fa1ddced9e
[RAG, Bart] Align RAG, Bart cache with T5 and other models of transformers ( #9098 )
...
* fix rag
* fix slow test
* fix past in bart
2020-12-14 12:32:26 +01:00
Patrick von Platen
9cc9f4122e
Make ProphetNetModel really compatible with EncoderDecoder ( #9033 )
...
* improve
* finish
* upload model
* fix lm head
* fix test
2020-12-11 16:59:54 +01:00
Patrick von Platen
06971ac4f9
[Bart] Refactor - fix issues, consistency with the library, naming ( #8900 )
...
* remove make on the fly linear embedding
* start refactor
* big first refactor
* save intermediate
* save intermediat
* correct mask issue
* save tests
* refactor padding masks
* make all tests pass
* further refactor
* make pegasus test pass
* fix bool if
* fix leftover tests
* continue
* bart renaming
* delete torchscript test hack
* fix imports in tests
* correct shift
* fix docs and repo cons
* re-add fix for FSTM
* typo in test
* fix typo
* fix another typo
* continue
* hot fix 2 for tf
* small fixes
* refactor types linting
* continue
* finish refactor
* fix import in tests
* better bart names
* further refactor and add test
* delete hack
* apply sylvains and lysandres commens
* small perf improv
* further perf improv
* improv perf
* fix typo
* make style
* small perf improv
2020-12-09 20:55:24 +01:00
Patrick von Platen
da37a21c89
push ( #9008 )
2020-12-09 15:14:33 +01:00
02d0e0355c
Diverse beam search 2 ( #9006 )
...
* diverse beam search
* bug fixes
* bug fixes
* bug fix
* separate out diverse_beam_search function
* separate out diverse_beam_search function
* bug fix
* improve code quality
* bug fix
* bug fix
* separate out diverse beam search scorer
* code format
* code format
* code format
* code format
* add test
* code format
* documentation changes
* code quality
* add slow integration tests
* more general name
* refactor into logits processor
* add test
* avoid too much copy paste
* refactor
* add to docs
* fix-copies
* bug fix
* Revert "bug fix"
This reverts commit c99eb5a8dc57a7b0d33a8ac06d8c6a32a7812ad4.
* improve comment
* implement sylvains feedback
Co-authored-by: Ayush Jain <a.jain@sprinklr.com >
Co-authored-by: ayushtiku5 <40797286+ayushtiku5@users.noreply.github.com >
2020-12-09 15:00:37 +01:00
Patrick von Platen
443f67e887
[PyTorch] Refactor Resize Token Embeddings ( #8880 )
...
* fix resize tokens
* correct mobile_bert
* move embedding fix into modeling_utils.py
* refactor
* fix lm head resize
* refactor
* break lines to make sylvain happy
* add news tests
* fix typo
* improve test
* skip bart-like for now
* check if base_model = get(...) is necessary
* clean files
* improve test
* fix tests
* revert style templates
* Update templates/adding_a_new_model/cookiecutter-template-{{cookiecutter.modelname}}/modeling_{{cookiecutter.lowercase_modelname}}.py
2020-12-02 19:19:50 +01:00
Patrick von Platen
5ced23dc84
[Pegasus] Refactor Tokenizer ( #8731 )
...
* refactor
* further refactor
* fix the rest tomorrow
* save intermediate
* finish slow tokenizer
* make more tests pass
* finish refactor
* fix comment
* clean further
* fix name
* fix naming
* Update src/transformers/models/reformer/tokenization_reformer.py
* Apply suggestions from code review
* Apply suggestions from code review
* refactor
* fix init tokenizers
* refactor
* improve convert
* refactor
* correct convert slow tokenizer
* final fix for Pegasus Tok
* remove ipdb
* improve links
2020-11-29 16:57:43 +01:00
Patrick von Platen
36b60ce9e8
fix mt5 config ( #8832 )
2020-11-28 19:50:49 +01:00
Patrick von Platen
a7d46a0609
Fix dpr<>bart config for RAG ( #8808 )
...
* correct dpr test and bert pos fault
* fix dpr bert config problem
* fix layoutlm
* add config to dpr as well
2020-11-27 16:26:45 +01:00
Patrick von Platen
a2cf37595e
[Flax test] Add require pytorch to flix flax test ( #8816 )
...
* try flax fix
* same for roberta
2020-11-27 14:40:42 +01:00
Patrick von Platen
8f07f5c44b
Revert "finetune.py: specifying generation min_length ( #8478 )" ( #8805 )
...
This reverts commit 5aa361f3e5 .
2020-11-26 20:12:01 +01:00
Patrick von Platen
2a6fbe6a40
[XLNet] Fix mems behavior ( #8567 )
...
* fix mems in xlnet
* fix use_mems
* fix use_mem_len
* fix use mems
* clean docs
* fix tf typo
* make xlnet tf for generation work
* fix tf test
* refactor use cache
* add use cache for missing models
* correct use_cache in generate
* correct use cache in tf generate
* fix tf
* correct getattr typo
* make sylvain happy
* change in docs as well
* do not apply to cookie cutter statements
* fix tf test
* make pytorch model fully backward compatible
2020-11-25 16:54:59 -05:00
Patrick von Platen
9c0afdaf7b
fix flaky ci ( #8694 )
2020-11-20 22:07:21 +01:00
Patrick von Platen
cdfa56afe0
[Tokenizer Doc] Improve tokenizer summary ( #8622 )
...
* improve summary
* small fixes
* cleaned line length
* correct "" formatting
* apply sylvains suggestions
2020-11-18 17:14:15 +01:00
Patrick von Platen
5104223552
[MT5] More docs ( #8589 )
...
* add docs
* make style
2020-11-17 12:47:57 +01:00
Patrick von Platen
86822a358b
T5 & mT5 ( #8552 )
...
* add mt5 and t5v1_1 model
* fix tests
* correct some imports
* add tf model
* finish tf t5
* improve examples
* fix copies
* clean doc
2020-11-17 12:23:09 +01:00
Patrick von Platen
f6cdafdec7
fix load weights ( #8528 )
...
* fix load weights
* delete line
2020-11-13 20:31:40 +01:00
Patrick von Platen
42e2d02e44
[T5] Bug correction & Refactor ( #8518 )
...
* fix bug
* T5 refactor
* refactor tf
* apply sylvains suggestions
2020-11-13 16:57:31 +01:00
Patrick von Platen
70708cca1a
fix t5 token type ids ( #8437 )
2020-11-10 14:21:54 -05:00
Patrick von Platen
b93569457f
fix t5 special tokens ( #8435 )
2020-11-10 18:54:17 +01:00
Patrick von Platen
9c83b96e62
[Tests] Add Common Test for Training + Fix a couple of bugs ( #8415 )
...
* add training tests
* correct longformer
* fix docs
* fix some tests
* fix some more train tests
* remove ipdb
* fix multiple edge case model training
* fix funnel and prophetnet
* clean gpt models
* undo renaming of albert
2020-11-09 18:24:41 +01:00
Patrick von Platen
07708793f2
fix encoder outputs ( #8368 )
2020-11-06 21:03:25 +01:00
Patrick von Platen
226b9debb7
Update PULL_REQUEST_TEMPLATE.md
2020-11-05 09:40:15 +01:00
Patrick von Platen
6f35c61f93
Update bug-report.md
2020-11-05 09:39:05 +01:00
Patrick von Platen
cb966e640b
[Generate Test] fix greedy generate test ( #8293 )
...
* fix greedy generate test
* delet ipdb
2020-11-04 15:44:36 +01:00
Patrick von Platen
068e6b5edd
make files independent ( #8267 )
2020-11-03 21:13:33 +01:00
Patrick von Platen
a1bbcf3f6c
Refactoring the generate() function ( #6949 )
...
* first draft
* show design proposition for new generate method
* up
* make better readable
* make first version
* gpt2 tests pass
* make beam search for gpt2 work
* add first encoder-decoder code
* delete typo
* make t5 work
* save indermediate
* make bart work with beam search
* finish beam search bart / t5
* add default kwargs
* make more tests pass
* fix no bad words sampler
* some fixes and tests for all distribution processors
* fix test
* fix rag slow tests
* merge to master
* add nograd to generate
* make all slow tests pass
* speed up generate
* fix edge case bug
* small fix
* correct typo
* add type hints and docstrings
* fix typos in tests
* add beam search tests
* add tests for beam scorer
* fix test rag
* finish beam search tests
* move generation tests in seperate file
* fix generation tests
* more tests
* add aggressive generation tests
* fix tests
* add gpt2 sample test
* add more docstring
* add more docs
* finish doc strings
* apply some more of sylvains and sams comments
* fix some typos
* make fix copies
* apply lysandres and sylvains comments
* final corrections on examples
* small fix for reformer
2020-11-03 16:04:22 +01:00
Patrick von Platen
9f1747f999
[Seq2Seq] Correct import in Seq2Seq Trainer ( #8254 )
2020-11-03 07:56:41 -05:00
Patrick von Platen
f744b81572
add new notebooks ( #8246 )
2020-11-02 20:21:55 +01:00
Patrick von Platen
dc26726df2
fix encoder decoder bug ( #8243 )
2020-11-02 20:12:34 +01:00
Patrick von Platen
5b178f3c87
Create README.md
2020-11-02 20:03:44 +01:00
Patrick von Platen
ebec410c71
Create README.md
2020-11-02 17:53:22 +01:00
Patrick von Platen
9bd30f7cf4
[Seq2SeqTrainer] Move import to init to make file self-contained ( #8194 )
...
* boom boom
* reverse order
2020-11-01 23:31:55 +01:00
Patrick von Platen
afa21504b1
add tags ( #8147 )
2020-10-29 12:45:55 +01:00
Patrick von Platen
664c7ec453
[Seq2Seq Trainer] Make sure padding is implemented for models without pad_token ( #8043 )
...
* make sure padding is implemented for non-padding tokens models as well
* add better error message
* add better warning
* remove results files
* Update examples/seq2seq/seq2seq_trainer.py
* remove unnecessary copy line
* correct usage of labels
* delete test files
2020-10-26 17:28:16 +01:00
Patrick von Platen
3c682ea15c
[Examples] Allow EncoderDecoderModels to be trained with Seq2Seq ( #7809 )
...
* Make Seq2Seq Trainer more similar to Trainer
* fix typo
* fix seq2seq trainer
* remove from tests
* remove lock
* remove train files
* delete test files
* correct typo
* check at init
* make sure trainer is not slowed down on TPU
* correct isort
* remove use cache
* fix use cache
* add last use chache = false
2020-10-23 23:05:51 +02:00
Patrick von Platen
4acfd1a8dc
[Reformer] remove reformer pad_token_id ( #7991 )
...
* remove reformer pad_token_id
* fix pegasus
2020-10-23 10:29:15 -04:00
Patrick von Platen and Sylvain Gugger
f34372a9ff
[PretrainedConfig] Fix save pretrained config for edge case ( #7943 )
...
* fix config save
* add test
* add config class variable and another test
* line break
* fix fsmt and typo
* god am I making many errors today :-/
* Update src/transformers/configuration_utils.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
2020-10-22 15:39:01 +02:00
Patrick von Platen
52decab371
fix test ( #7947 )
2020-10-21 19:06:23 +02:00
Patrick von Platen
9b6610f7f6
[ProphetNet] Correct Doc string example ( #7944 )
...
* correct xlm prophetnet auto model and examples
* fix line-break docs
2020-10-21 17:27:20 +02:00
Patrick von Platen
5cd9e2cba1
Update README.md
2020-10-21 12:43:42 +02:00
Patrick von Platen
220b5f97ca
Create README.md
2020-10-21 12:34:46 +02:00
Patrick von Platen
8ffd7fb12d
Update README.md
2020-10-21 12:27:09 +02:00
Patrick von Platen
613ab364eb
Update README.md
2020-10-21 12:23:17 +02:00
Patrick von Platen
f7eb17dc47
Update README.md
2020-10-21 12:19:44 +02:00
Patrick von Platen
29792864cb
[ProphetNet] Add Question Generation Model + Test ( #7942 )
...
* new prophetnet model
* correct name
* make style
2020-10-21 11:49:58 +02:00
Patrick von Platen
0264048660
Update README.md
2020-10-20 16:13:49 +02:00
Patrick von Platen
ffd675b42c
add summary ( #7927 )
2020-10-20 10:11:02 -04:00
Patrick von Platen
f3312515b7
Add note for WikiSplit
2020-10-20 15:42:29 +02:00
Patrick von Platen
0724c0f3a2
Fix EncoderDecoder WikiSplit Example
2020-10-20 15:13:22 +02:00
Patrick von Platen
c912ba5f69
[EncoderDecoder] Fix Typo ( #7915 )
...
* fix encoder decoder models
* add .gitignore
2020-10-19 22:02:42 +02:00
Patrick von Platen
e3d2bee8d0
fix t5 training docstring ( #7911 )
2020-10-19 21:49:47 +02:00
Patrick von Platen
f5c45a19e6
Fix Rag example docstring ( #7872 )
...
* fix rag examples
* fix token generate example
2020-10-17 22:46:47 +02:00
Patrick von Platen
dc552b9b70
Fix typo in sequence model card
2020-10-16 16:05:06 +02:00
Patrick von Platen and Thomas Wolf
82b09a8481
[Rag] Fix loading of pretrained Rag Tokenizer ( #7756 )
...
* fix rag
* Update tokenizer save_pretrained
Co-authored-by: Thomas Wolf <thomwolf@users.noreply.github.com >
2020-10-13 14:34:22 +02:00
Patrick von Platen
2d4e928d97
Update PULL_REQUEST_TEMPLATE.md
...
Putting my name on a couple more issues to directly redirect them to me
2020-10-13 12:18:31 +02:00
Patrick von Platen
bd2621583b
fix data type ( #7513 )
2020-10-01 18:15:41 +02:00
Patrick von Platen and Sylvain Gugger
62f5ae68ec
[Seq2Seq] Fix a couple of bugs and clean examples ( #7474 )
...
* clean T5
* fix t5 tests
* fix index typo
* fix tf common test
* fix examples
* change positional ordering for Bart and FSTM
* add signature test
* clean docs and add tests
* add docs to encoder decoder
* clean docs
* correct two doc strings
* remove sig test for TF Elektra & Funnel
* fix tf t5 slow tests
* fix input_ids to inputs in tf
* Update src/transformers/modeling_bart.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
* Update src/transformers/modeling_bart.py
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
* implement lysandre results
* make style
* fix encoder decoder typo
* fix tf slow tests
* fix slow tests
* renaming
* remove unused input
Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com >
2020-10-01 17:38:50 +02:00
Patrick von Platen
8279471506
correct RAG model cards ( #7420 )
2020-09-28 11:08:39 +02:00
Patrick von Platen
e50a931c11
[Longformer, Bert, Roberta, ...] Fix multi gpu training ( #7272 )
...
* fix multi-gpu
* fix longformer
* force to delete unnecessary layers
* fix notifications
* fix warning
* fix roberta
* fix tests
* remove hasattr
* fix tests
* fix roberta
* merge and clean authorized keys
2020-09-25 20:33:21 +02:00
Patrick von Platen
2c8ecdf8a8
fix rag retriever save pretrained ( #7399 )
2020-09-25 19:47:12 +02:00
Patrick von Platen
1a14687e6f
Update README.md
2020-09-25 19:43:48 +02:00
Patrick von Platen
3327c2b0f6
Update README.md
2020-09-25 19:43:36 +02:00
Patrick von Platen
4e5b036bdd
Update README.md
2020-09-25 18:16:46 +02:00
Patrick von Platen
55eccfbb49
Update README.md
2020-09-25 18:16:44 +02:00
Patrick von Platen
5ff0d6d7d0
Update README.md
2020-09-25 16:58:29 +02:00
Patrick von Platen
571c7a11c1
[Rag] Fix wrong usage of num_beams and bos_token_id in Rag Sequence generation ( #7386 )
...
* fix_rag_sequence
* add second bug fix
2020-09-25 14:35:49 +02:00
Patrick von Platen
2dd652d757
[RAG] Add missing doc and attention_mask to rag ( #7382 )
...
* add docs
* add missing docs and attention_mask in fine-tune
2020-09-25 11:23:55 +02:00
Patrick von Platen
0804d077c6
correct attention mask ( #7373 )
2020-09-24 23:22:04 +02:00
Patrick von Platen
0cbe1139b1
Update README.md
2020-09-21 11:53:08 +02:00
Patrick von Platen
9397436ea5
Create README.md
2020-09-18 16:52:00 +02:00
Patrick von Platen
7eeca4d399
Create README.md
2020-09-18 16:44:02 +02:00
Patrick von Platen
31516c776a
Update README.md
2020-09-18 16:37:14 +02:00
Patrick von Platen
4c14669a78
Update README.md
2020-09-18 16:35:11 +02:00
Patrick von Platen
afd6a9f827
Create README.md
2020-09-18 11:41:12 +02:00
Patrick von Platen
9f1544b9e0
Create README.md
2020-09-18 11:37:20 +02:00
Patrick von Platen
85ffda96fc
fix encoder decoder kwargs ( #7131 )
2020-09-15 21:10:07 +02:00
Patrick von Platen
7af2791d77
Create README.md
2020-09-15 16:47:36 +02:00
Patrick von Platen
221d4c63a3
clean naming ( #7068 )
2020-09-11 09:57:53 +02:00
Patrick von Platen
db38f7ce29
[BertGeneration, Docs] Fix another old name in docs ( #7050 )
...
* correct docs for bert generation
* upload
2020-09-10 17:12:33 +02:00
Patrick von Platen
3bd95b0faf
correct docs for bert generation ( #7048 )
2020-09-10 17:08:40 +02:00
Patrick von Platen
eb2feb5d90
Create README.md
2020-09-10 17:05:50 +02:00
Patrick von Platen
9ccdb1d517
Update README.md
2020-09-10 17:01:19 +02:00
Patrick von Platen
60698936fc
Create README.md
2020-09-10 17:00:10 +02:00
Patrick von Platen
e0c3bc8ee0
Create README.md
2020-09-10 16:51:15 +02:00
Patrick von Platen
c356b9878d
Create README.md
2020-09-10 16:45:44 +02:00
Patrick von Platen
5afd3f6196
Create README.md
2020-09-10 16:44:47 +02:00
Patrick von Platen
7fd1febf38
Add "Leveraging Pretrained Checkpoints for Generation" Seq2Seq models. ( #6594 )
...
* add conversion script
* improve conversion script
* make style
* add tryout files
* fix
* update
* add causal bert
* better names
* add tokenizer file as well
* finish causal_bert
* fix small bugs
* improve generate
* change naming
* renaming
* renaming
* renaming
* remove leftover files
* clean files
* add fix tokenizer
* finalize
* correct slow test
* update docs
* small fixes
* fix link
* adapt check repo
* apply sams and sylvains recommendations
* fix import
* implement Lysandres recommendations
* fix logger warn
2020-09-10 16:40:51 +02:00
Patrick von Platen
63e539459d
Update README.md
2020-09-10 16:34:28 +02:00