Stuck on an issue?

Lightrun Answers was designed to reduce the constant googling that comes with debugging 3rd party libraries. It collects links to all the places you might be looking at while hunting down a tough bug.

And, if you’re still stuck at the end, we’re happy to hop on a call to see how we can help out.

Predictions for certain paragraphs are inaccurate. What's wrong?

See original GitHub issue

I have been testing cdQA on paragraphs that have been generated from CSV. I convert the structured data into English, then predict answers using BERT.

I’ve described the approach here: https://datascience.stackexchange.com/questions/58186/transform-data-into-english-then-predict-an-answer-using-bert

I combine 2 or 3 sentences into paragraphs, then concatenate multiple paragraphs into one dataframe for cdQA pipeline, then query the dataset but results are often incorrect. An example of a sentence:

According to our website, the Melbourne Convention Centre & South Wharf Precinct
project is located at 1 Convention Centre Pl, South Wharf VIC 3006, Australia. The
Melbourne Convention Centre & South Wharf Precinct project has won three awards. The project started in 2014 and was completed in 2016.

And query

how many awards has the Melbourne Convention Centre project won?

Could this form of English writing be too dissimilar to the corpora and datasets on which BERT was pre-trained and fine-tuned? Can you suggest how I could improve results? Thanks.

Issue Analytics

State:
Created 4 years ago
Comments:25

Top GitHub Comments

1reaction

radsimucommented, Sep 24, 2019

finally I found the problem: readme.md says cdqa_pipeline = QAPipeline(model='bert_qa_vCPU-sklearn.joblib')

should be cdqa_pipeline = QAPipeline(reader='bert_qa_vCPU-sklearn.joblib')

1reaction

JimAvacommented, Sep 19, 2019

Hi @andrelmfarias ,

I’ve been trying out the newly introduced retrieve based on BM25 and it seems to be working great! Thank you for your efforts. Will provide more updates…