Downloads · 30 days
37
100% of all-time downloads
khursheed/datacard-ci
datacard-ci is a text generation model from khursheed. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
Your dataset README says the IDs are unique. Are they still unique after the next update?
Downloads · 30 days
37
100% of all-time downloads
All-time downloads
37
Public
Repo size
33.7 MB
Likes
0
Public
Click a slice to open those files.
.zip13.1 MB · 41%
From the Hugging Face model README
Your dataset README says the IDs are unique. Are they still unique after the next update?
I built DataCard CI to help check promises like that. The model reads a short claim and the table schema, then suggests a check. You review what it meant before the verifier runs against your CSV files.
The latest release, v0.4, makes that review-and-check workflow usable without loading a model. It runs on CPU, works offline after installation, and needs no API key. The Qwen3-0.6B LoRA adapter can help draft checks; each suggestion needs review before use.
NA.The model proposes the first three operations. Missing-value checks are available in the verifier when you write the check explicitly.
A result stays tied to the documentation, the reviewed check, the exact CSV files and the verifier version. Change any of those inputs and the workflow asks for a new review. It won't silently reuse an old approval on new data.
Download the v0.4 quick-start notebook and open it through Colab's File → Upload notebook menu. Choose a CPU runtime. It installs the small verifier wheel, runs a toy example, then changes the data so you can see the old review get rejected. No training or Hugging Face login is required.
For a local installation, download the wheel and run:
python3 -m pip install --no-index datacard_ci-0.4.0-py3-none-any.whl
datacard-review --help
The source archive includes the tests, examples and workflow instructions. Python 3.10 or newer is required. The isolated installation test also needs host pip 22.3 or newer.
I tested the release locally and in a real Google Colab CPU session:
ensurepip. I fixed the runner and repeated the check successfully.The release also rejects malformed JSON, invalid UTF-8, unsupported checks and blank physical CSV records. It limits file size and total cells instead of quietly checking only part of a large file. See the Colab report, raw test evidence and release notes.
The training dataset is available separately. For evaluation details and supported-use boundaries, see READINESS.md. The test counts above cover the software and review workflow.
The adapter needs the Qwen3-0.6B base model plus PyTorch, Transformers and PEFT. The research helper predict_v2.py returns an unapproved proposal and rejects unsupported output. For the model-loading example and pinned settings, see the adapter usage guide. Use the new CPU quick-start notebook for the v0.4 verifier workflow.
This repository supplies files, not a hosted inference service. Downloading them does not use my API credentials. If you choose a paid runtime or inference provider, its compute is your own choice and cost. There is no need to start a GPU for the verifier.
This checks small, complete CSV files: up to 10 splits, 2 MB and 20,000 rows per file, 100 columns and 500,000 data cells across the batch. It accepts up to 100 claims and 64 KB each for documentation and claims JSON. Every nonzero checking-command exit should block a pipeline.
A pass applies to the selected checks and supplied files. It does not establish consent, redistribution rights, fairness or overall dataset quality. Review records are local attestations, not authenticated signatures. Reports include documentation text, so keep confidential reports out of public logs. More detail is in SECURITY.md.
If dataset checks are a recurring pain point, start a Hugging Face Discussion. Tell me what you check by hand, what tends to break, and what you'd want caught before training or publishing.
A small invented example is enough: a claim, the column names, the check you expected, and what happened. Critique is welcome—including “I'd rather write this check myself.” Please leave out private records and credentials.
Original project code and synthetic examples are MIT licensed. The Qwen base is Apache-2.0; no base weights are included. The older UCI Wine Quality integration fixture has separate CC BY 4.0 attribution. This v0.4 distribution contains original code and toy examples, not third-party research datasets.