Add token-level recall metric (get_metric_recall_default_prep) - #4
Merged
Merged
Conversation
Adds compare_recall and a get_metric_recall_default_prep() factory mirroring the existing F1 metric, exposing the recall component of the SQuAD-style token overlap. Useful for report/VQA evaluation where recall is reported alongside F1 and accuracy.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add a token-level recall metric
This adds
get_metric_recall_default_prep()and acompare_recallcomparisonfunction to
ovqa/metrics/simple.py, mirroring the existingget_f1_score_default_prep()/compare_f1.compare_recallreusestransformers' SQuADget_tokens(already used bycompute_f1) and returns the recall component of the token overlap — thefraction of reference tokens that also appear in the candidate — with the same
empty-string edge-case handling as the SQuAD F1.
Why
Report / VQA evaluations often report recall alongside F1 and accuracy, but OVQA
only exposes F1 and exact/contains comparisons. This is a small, dependency-free
addition consistent with the existing metric factories.
Notes
collectionsis stdlib;get_tokenscomes from the sametransformers.data.metrics.squad_metricsmodule already imported for F1).