Stefan Schweter
19fa01ce2a
token-classification: use is_world_process_zero instead of deprecated is_world_master() ( #8828 )
2020-11-30 09:21:56 -05:00
Stefan Schweter
185259c261
[model_cards] Update Italian BERT models and introduce new Italian XXL ELECTRA model 🎉 ( #8343 )
2020-11-06 03:17:03 -05:00
Stefan Schweter
cfa26d2b41
github: add @stefan-it to bug-report template for all token-classification related bugs ( #6489 )
2020-08-18 08:38:54 -04:00
Stefan Schweter
d812e6d76e
NER: fix construction of input examples for RoBERTa ( #4943 )
...
* utils_ner: do not add extra sep token for RoBERTa model
* run_pl_ner: do not add extra sep token for RoBERTa model
2020-06-15 08:30:40 -04:00
Stefan Schweter
2a4b9e09c0
NER: Add new WNUT’17 example ( #4681 )
...
* ner: add preprocessing script for examples that splits longer sentences
* ner: example shell scripts use local preprocessing now
* ner: add new example section for WNUT’17 NER task. Remove old English CoNLL-03 results
* ner: satisfy black and isort
2020-06-04 19:13:17 -04:00
Stefan Schweter
15d45211f7
[model_cards]: 🇹🇷 Add new ELECTRA small and base models for Turkish ( #4318 )
2020-05-12 15:01:17 -04:00
Stefan Schweter
3f42eb979f
Documentation: fix links to NER examples ( #4279 )
...
* docs: fix link to token classification (NER) example
* examples: fix links to NER scripts
2020-05-11 12:48:21 -04:00
Stefan Schweter
1e616c0af3
NER: parse args from .args file or JSON ( #4110 )
...
* ner: parse args from .args file or JSON
* examples: mention json-based configuration file support for run_ner script
2020-05-02 10:29:17 -04:00
Stefan Schweter
e80be7f1d0
docs: add xlm-roberta section to multi-lingual section ( #4101 )
2020-05-01 11:06:58 -04:00
Stefan Schweter
b5c6d3d4c7
notebooks: minor fix for community provided models example ( #4025 )
2020-04-28 09:12:25 +02:00
Stefan Schweter
601ac5b1dc
[model_cards]: use MIT license for all dbmdz models
2020-03-27 18:06:25 -04:00
Stefan Schweter
b31ef225cf
[model_cards] 🇹🇷 Add new (uncased, 128k) BERTurk model
2020-03-24 11:29:06 -04:00
Stefan Schweter
b4009cb001
[model_cards] 🇹🇷 Add new (cased, 128k) BERTurk model
2020-03-24 11:29:06 -04:00
Stefan Schweter
d3283490ef
[model_cards] 🇹🇷 Add new (uncased) BERTurk model
2020-03-24 11:29:06 -04:00
Stefan Schweter
14e455b716
[model_cards] 🇹🇷 Add new (cased) DistilBERTurk model
2020-03-11 18:40:38 -04:00
Stefan Schweter
c88ed74ccf
[model_cards] 🇹🇷 Add new (cased) BERTurk model
2020-02-17 09:54:46 -05:00
Stefan Schweter
3376adc051
configuration/modeling/tokenization: add various fine-tuned XLM-RoBERTa models for English, German, Spanish and Dutch (CoNLL datasets)
2019-12-19 21:30:23 +01:00
Stefan Schweter
a26ce4dee1
examples: add XLM-RoBERTa to glue script
2019-12-19 02:23:01 +01:00
Stefan Schweter
fe9aab1055
tokenization: use S3 location for XLM-RoBERTa model
2019-12-18 23:47:48 +01:00
Stefan Schweter
5c5f67a256
modeling: use S3 location for XLM-RoBERTa model
2019-12-18 23:47:00 +01:00
Stefan Schweter
db90e12114
configuration: use S3 location for XLM-RoBERTa model
2019-12-18 23:46:33 +01:00
Stefan Schweter
f09d999641
docs: fix numbering 😅
2019-12-18 19:49:33 +01:00
Stefan Schweter
dd7a958fd6
docs: add XLM-RoBERTa to pretrained model list (incl. all parameters)
2019-12-18 19:45:46 +01:00
Stefan Schweter
d35405b7a3
docs: add XLM-RoBERTa to index page
2019-12-18 19:45:10 +01:00
Stefan Schweter
3e89fca543
readme: add XLM-RoBERTa to model architecture list
2019-12-18 19:44:23 +01:00
Stefan Schweter
128cfdee9b
tokenization add XLM-RoBERTa base model
2019-12-18 19:28:16 +01:00
Stefan Schweter
e778dd854d
modeling: add XLM-RoBERTa base model
2019-12-18 19:27:34 +01:00
Stefan Schweter
64a971a915
auto: add XLM-RoBERTa to auto tokenization
2019-12-18 18:24:32 +01:00
Stefan Schweter
036831e279
auto: add XLM-RoBERTa to audo modeling
2019-12-18 18:23:42 +01:00
Stefan Schweter
41a13a6375
auto: add XLMRoBERTa to auto configuration
2019-12-18 18:20:27 +01:00
Stefan Schweter
01b68be34f
converter: remove XLM-RoBERTa specific script (can be done with the script for RoBERTa now)
2019-12-18 12:24:46 +01:00
Stefan Schweter
ca31abc6d6
tokenization: *align* fairseq and spm vocab to fix some tokenization errors
2019-12-18 11:36:54 +01:00
Stefan Schweter
cce3089b65
Merge remote-tracking branch 'upstream/master' into xlmr
2019-12-18 11:05:16 +01:00
Stefan Schweter
8c276b9c92
Merge branch 'master' into distilbert-german
2019-11-27 18:11:49 +01:00
Stefan Schweter
da06afafc8
tree-wide: add trailing comma in configuration maps
2019-11-19 21:57:00 +01:00
Stefan Schweter
2e2c0375c3
distilbert: add German distilbert model to positional embedding sizes map
2019-11-19 20:41:18 +01:00
Stefan Schweter
e7cf2ccd15
distillation: add German distilbert model
2019-11-19 19:55:19 +01:00
Stefan Schweter
e631383d4f
docs: add new German distilbert model to pretrained models
2019-11-19 19:52:40 +01:00
Stefan Schweter
f21dfe36ba
distilbert: add vocab for new German distilbert model
2019-11-19 19:51:31 +01:00
Stefan Schweter
22333945fb
distilbert: add pytorch model for new German distilbert model
2019-11-19 19:51:01 +01:00
Stefan Schweter
337802783f
distilbert: add configuration for new German distilbert model
2019-11-19 19:50:32 +01:00
Stefan Schweter
a1c34bd286
distillation: fix ModuleNotFoundError error in token counts script
2019-08-31 12:21:38 +02:00
Stefan Schweter
e6cc6d237f
docs: fix link to various notebooks
2019-07-16 23:42:28 +02:00
Stefan Schweter
5b78400e21
docs: fix link to modeling example source (bert)
2019-07-16 23:41:57 +02:00
Stefan Schweter
61cc3ee350
docs: fix link to tf checkpoint to pytorch script
2019-07-16 23:41:04 +02:00
Stefan Schweter
dbbd94cb7a
docs: fix link to bertology example and update dataset description
2019-07-16 23:40:04 +02:00