Downloads · 30 days
29
100% of all-time downloads
Atharv-kumar/visual-question-answering
visual-question-answering is a visual question answering model from Atharv-kumar. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
A structured set of research notes on Visual Question Answering, with concrete evaluation references and open questions. Plans and hypotheses are kept separate from completed results.
Downloads · 30 days
29
100% of all-time downloads
All-time downloads
29
Public
Parameters
49.6K
199 KB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors199 KB · 97%
From the Hugging Face model README
A structured set of research notes on Visual Question Answering, with concrete evaluation references and open questions. Plans and hypotheses are kept separate from completed results.
Start with paper_notes.md for the full note. Sections labeled as plans or hypotheses should not be interpreted as experimental results. If results are added later, they should include dataset versions, commands, seeds, hardware, and raw logs.
The note is intentionally exploratory. It does not claim benchmark improvements, completed ablations, released code, or a trained checkpoint. References and proposed datasets provide a starting point for verification rather than evidence that the study has already been run.
paper_notes.md — primary artifactREADME.md — this documentationReleased under cc-by-4.0. Review the source-data terms separately when this repository is used with external datasets.