HSH·Intelligence
Data-on-Demand · for AI fine-tuning

The dataset is the half of fine-tuning nobody hands you.

Describe the task. We build a clean, answer-verified dataset and deliver it HuggingFace-ready — drop the repo straight into Gradients, TRL, Axolotl, or Unsloth. Every checkable answer is verified in code, not trusted from a model.

row 0428 · verified-math-reasoning-3ktrain split
instructionAn item costs $161. It is on sale with a 25% discount. What is the final price in dollars?
output…25% of $161 = $40.25 · $161 − $40.25 = $120.75 · The answer is 120.75
ground truth120.75  computed in Python, not the model
answer matches ground truth — row kept · mismatches are discarded
How it works

Python owns the truth. The model only writes the prose.

Most synthetic datasets trust the language model to be right. This one doesn't. The correct answer is computed independently before the model writes a word — then the model's answer is checked against it.

step 01

Construct

Each problem is built in code with a known-correct answer as ground truth.

step 02

Generate

The model writes step-by-step reasoning and a final answer for the problem.

step 03

Verify

The model's answer is checked against ground truth in code. Mismatches are discarded.

step 04

Deliver

Deduplicated, split train/val/test, documented, pushed to a HuggingFace repo.

100%
of checkable answers verified against code-computed ground truth
Alpaca
instruction / input / output — Gradients-ready, drop-in for TRL · Axolotl · Unsloth
24h
standard turnaround on a verifiable build · Apache-2.0, commercial use
Building this into an autonomous agent? The same verified-dataset engine is callable programmatically over MCP — your agent describes the task, pays per-tier in USDC via x402, and receives the HuggingFace repo. No card, no form.
Agent & developer docs →
Order

Tell us what you need. Get a price in seconds.

Describe your dataset and we'll price it instantly for standard work, or quote it within 24h for specialized domains. Every order gets a scope review — if we can't build it as described, you get a full refund.

Your live price
$75
Tier S · Standard
Updates as you choose
Standard tasks are instant-buy. We build it and email your HuggingFace repo in 3–5 days.
Agents: pay via x402 →
Building this into an autonomous agent? Agents discover and buy this dataset programmatically over MCP — describe the task, get a quote, pay per-call in USDC via x402. No card, no form.
Agent & developer docs →
Pricing

Priced like a fine-tuning job — not a data subscription.

A Gradients run costs $100–500 and you still have to bring the data. We supply the verified dataset for a flat per-build price. No setup fee, no per-record meter.

Tier S
$75
1,000 – 2,000 rows
For indie devs & first-time fine-tuners testing a hypothesis.
Narrow-task LoRA — format compliance, single-skill injection, or style alignment. Teaches the model the shape of one task.
  • Answer-verified rows
  • Train / val / test split
  • Dataset card + license
  • HuggingFace repo
Best for: "output this JSON," "classify into these 4 buckets," "reply in this voice."
Tier M · most fine-tunes
$150
2,000 – 5,000 rows
For startups & ML engineers shipping to staging.
LoRA on a domain or a small full fine-tune. The model learns a skill set — varied inputs within a domain, internal tools, API-calling behavior.
  • Everything in S
  • Larger, richer coverage
  • Schema tuned to your trainer
  • Ready to drop into any fine-tuning tool
Best for: domain chatbots, internal-API code models, multi-skill assistants.
Tier L
$300
5,000 – 10,000 rows
For teams going to production who need a domain model they can trust.
Full fine-tune territory — genuine domain expertise and multi-task capability. The model stops parroting and starts handling ambiguous inputs reliably.
  • Everything in M
  • Priority build queue (2–3 days)
  • Custom schema & task review
  • Adversarial + edge-case test set
Best for: legal analysis, clinical summarization, specialized coding assistants.
Tier XL
$600
15,000 – 20,000 rows
For AI labs & serious domain-model builders who need depth.
Serious domain specialization — multi-task training, agent behavior shaping, deep single-domain expertise. Build a domain specialist, not a tuned generalist.
  • Everything in L
  • Multi-task schema + task routing
  • Instruction diversity at scale
  • Async delivery (5–7 days)
Best for: multi-turn domain workflows, in-domain tool use, cross-task reasoning.
Custom
Quote
25,000+ rows / bespoke
For anyone whose scope doesn't fit a tier — nobody left out. Multi-language coverage, adversarial red-team sets, synthetic augmentation, multi-modal, or very large corpora. We scope it with you and quote a flat price — tell us the task, domain, and row target, and we quote within 24h.
Scope & guarantee

Every order is reviewed. If we can't build it, you don't pay.

24-hour scope review

After payment, every order gets a scope review within 24 hours. If your request needs domain expertise, sources, or verification depth beyond the quoted price, we propose an adjusted scope first — reduce rows, simplify verification, or refocus the task. If you accept, we build immediately.

Full-refund fallback

If no acceptable adjustment exists, you get a full refund within 2 business days — no partial charges, no quibbling. Revisions for errors in the delivered dataset (wrong format, bad split, missing rows) are always included. Changes to the task definition itself are a new order.

What we don't build

To keep every promise we make, we don't fulfill requests requiring: HIPAA-regulated patient data, classified or restricted-source material, real-time data streams, or expert credentials we don't hold (e.g. board-certified clinical sign-off). These are flagged at scope review and automatically refunded.