haystack/test/benchmarks/reader_results.csv
Branden Chan 1cebcb7dda
Create time and performance benchmarks for all readers and retrievers (#339)
* add time and perf benchmark for es

* Add retriever benchmarking

* Add Reader benchmarking

* add nq to squad conversion

* add conversion stats

* clean benchmarks

* Add link to dataset

* Update imports

* add first support for neg psgs

* Refactor test

* set max_seq_len

* cleanup benchmark

* begin retriever speed benchmarking

* Add support for retriever query index benchmarking

* improve reader eval, retriever speed benchmarking

* improve retriever speed benchmarking

* Add retriever accuracy benchmark

* Add neg doc shuffling

* Add top_n

* 3x speedup of SQL. add postgres docker run. make shuffle neg a param. add more logging

* Add models to sweep

* add option for faiss index type

* remove unneeded line

* change faiss to faiss_flat

* begin automatic benchmark script

* remove existing postgres docker for benchmarking

* Add data processing scripts

* Remove shuffle in script bc data already shuffled

* switch hnsw setup from 256 to 128

* change es similarity to dot product by default

* Error includes stack trace

* Change ES default timeout

* remove delete_docs() from timing for indexing

* Add support for website export

* update website on push to benchmarks

* add complete benchmarks results

* new json format

* removed NaN as is not a valid json token

* fix benchmarking for faiss hnsw queries. do sql calls in update_embeddings() as batches

* update benchmarks for hnsw 128,20,80

* don't delete full index in delete_all_documents()

* update texts for charts

* update recall column for retriever

* change scale and add units to desc

* add units to legend

* add axis titles. update desc

* add html tags

Co-authored-by: deepset <deepset@Crenolape.localdomain>
Co-authored-by: Malte Pietsch <malte.pietsch@deepset.ai>
Co-authored-by: PiffPaffM <markuspaff.mp@gmail.com>
2020-10-12 13:34:42 +02:00

864 B

1EMf1top_n_accuracytop_nreader_timeseconds_per_querypassages_per_secondreadererror
200.75897522332715320.80679857946718850.96713298499915725133.797060279999980.01127566663408056492.30397120949361deepset/roberta-base-squad2
310.73596831282656330.78233062653186860.97143097926849825125.223233931999970.01055311258486431798.62387044489225deepset/minilm-uncased-squad2
420.7008258890948930.74902716000535050.95853699646047535123.589592784999920.01041543846157086799.92750782409666deepset/bert-base-cased-squad2
530.78215068262261920.82645457080974720.97623461992246755312.422336850999950.02632920418430810239.529824033964466deepset/bert-large-uncased-whole-word-masking-squad2
640.80996123377717850.85262751909545860.97724591269172425314.31798548199980.02648895883043989739.29142006004379deepset/xlm-roberta-large-squad2