Hugging Face Blog: Writer Used ML Intern Agent to Build Seven Small Models From Prompts
A Hugging Face account published October 8, 2026 describes using ML Intern, an agent inside HuggingChat, to build seven small models from written prompts, each ending as a public model on the Hub with evaluation in its model card. Reported compute costs totalled about USD 103, itemised as GPU and CPU job charges, with individual projects ranging from about USD 1.90 to about USD 37.
David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.
Executive Summary
- Hugging Face published an account describing how a writer used ML Intern, an agent in HuggingChat, to build seven small models from prompts, each ending as a public model on the Hub with evaluation in its model card (source).
- The first project produced a 0.8B prompt rewriter that returns valid output 99.7% of the time and runs on a CPU, using about a quarter of the tokens of the 9B teacher (source).
- According to Hugging Face's blog, compute costs across the projects totalled about USD 103, itemised by GPU and CPU job charges for each session, with individual projects ranging from about USD 1.90 to about USD 37 (source).
- ML Intern is described as planning work, requesting a budget before spending, running a smoke test before the full job, then training, evaluating and publishing on Hugging Face hardware (source).
Key Takeaways
- The reported workflow replaces manual fine-tuning and data-pipeline work with a prompt that names the dataset, the base model and the training script.
- Half of the seven projects are described as filling gaps the author found on the Hub, including a camera-angle LoRA and a doodle-in LoRA for Qwen-Image 2.1.
- The author attributes the results to prompt structure rather than model novelty, citing two required instructions: a zero-shot baseline before training and a smoke test with an attached check.
- All cost figures are job charges reported for each session and reflect compute only, not the value of the author's time or the underlying research.
Hugging Face account describes an agentic build path for seven small models
The article's central claim is operational rather than architectural. A writer identified as yuvraj sharma, posting alongside Abubakar Abid, describes wanting a small version of the prompt rewriter shipped with Qwen-Image 2.1. The official rewriter is a 9B model that the author says needs about 20 GB of memory and reasons for thousands of tokens before producing a paragraph. On the Hub, the author found only compressed copies of the same 9B model.
The response was to describe the target model to ML Intern. The next day, the author reports, a 0.8B version existed that runs on a CPU, returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. Compute for the whole project, including having the 9B model label 8,797 example requests, came to USD 16.
Five further models followed the same pattern: each began as a message in HuggingChat with ML Intern switched on, and each ended as a public model on the Hub with its evaluation recorded in the model card. The claim worth noting for buyers is not that small models are good, but that the surrounding work — dataset construction, baseline measurement, checkpoint selection, evaluation and publication — was delegated to an agent inside an existing chat surface.
Hugging Face prompt design is described as the main lever
The author states that the first message is where the effort goes. The initial prompt for the citrus model ran about 450 words; by the sixth project it was closer to 2,000, because each project taught the author something he wanted carried into the next. All seven prompts are published on GitHub at yvrjsharma/ml-intern-prompts, presented exactly as written.
The described structure is consistent across projects. A prompt opens with the idea in one line and the reason for it, then names the exact pieces: the dataset, the base model and the training script. Material the author has already checked goes under a heading that reads "Verified facts, do not re-derive", so the agent spends its budget on work rather than rediscovering known information. For the camera-angle LoRA, that section recorded which trainer had just added transparent-image support and which open GitHub issues made the fallback trainer risky.
Two instructions are described as critical. The first asks for a baseline before any training; the citrus prompt reportedly says, "Also report the base model's zero-shot score on the same metric before training so we can see the gain." The author's stated consequence of omitting it is a trained model with no evidence that it improves on the starting point. The second is a smoke test with a check attached — for image LoRAs, 50 training steps followed by verification that saved weights actually changed, before paying for the full run.
Related: AI Film Making Startups 2026: Top Players and Market Trajectories
Budget control is described as structural rather than advisory. According to Hugging Face's blog, ML Intern begins every task with a zero dollar budget and needs permission before executing paid jobs, so a cap such as "Cap total spend at USD 12 and ask me before exceeding it" is enforced by the agent's own gating. When no budget is given, the agent reportedly proposes a couple of paths depending on project size and asks which one is preferred.
Hugging Face projects span vision, image generation and distillation
The citrus-disease-vlm-instruct dataset, merged from three sources hosted by the Project-AgML organization on the Hub, contains 3,017 annotated images across 21 distinct pests, illnesses, nutritional gaps and treatment approaches. ML Intern fine-tuned Qwen3.5-2B on those examples and benchmarked the foundation model first. On 335 test photos, the base model named the right problem 14.9% of the time; after two epochs on one A10G, the fine-tuned model reached 52.8%. Compute cost was about USD 1.90.
The Huggy LoRA targeted FLUX.2 klein base 4B, trained on 84 captioned drawings from the Chunte/huggy_for_training dataset. The agent saved a checkpoint every 100 steps and drew the same prompts with each, which the author says made selection easy. Step 200 was the first where the character was fully on-model; from step 500 onward, the author reports the style bled into unrelated prompts. The LoRA also works on the distilled klein model at 4 steps. Compute cost was about USD 7.60.
For deeper context, see our AI analysis: "Microsoft Azure AI Platform Leader in Gartner Cloud Ranking".
The Viewpoint Orbit LoRA was built because, days after the Qwen-Image 2.1 release, the author found no camera-angle LoRA for it. ML Intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each — 24,722 transparent images — in a CPU job costing a few cents, then finalised 461 objects for training and 40 held out for testing, with 1,844 before-and-after pairs spread evenly over 23 camera instructions. Training ran 2,000 steps in about 90 minutes on one A100, roughly USD 3.75. The project took about half a day and 48 jobs, including failures on missing packages or wrong paths that ML Intern resubmitted. Total compute cost was about USD 16.
The Doodle-in LoRA had no existing dataset, so the prompt described how to build one: start from a real photo in Open Images, remove one object with the LaMa inpainting model, draw a scribble where the object had been, and treat the untouched photo as the target. ML Intern wrote and tested the pair-building scripts in a CPU sandbox, then ran them as GPU jobs while recording the author and licence of every source photo. It built 6,042 training pairs and a 160-pair test set, 40 of which came from 23 object classes kept out of training entirely. Training ran 2,000 steps in 1 hour 38 minutes on one A100, about USD 4, and a comparison of saved checkpoints on 48 test pairs selected step 500. Paired with the Viggle turbo LoRA at 6 steps, 67.5% of objects were detected where drawn, at 4.7 seconds per edit, and objects from the 23 unseen classes landed at 65.0% versus 64.2% for the rest. The project took a little over a day and 59 jobs, at about USD 24.
Two distillation efforts are also described. For the Pocket Rewriter, ML Intern generated 8,797 short image requests with a small instruct model through Inference Providers, following a mix set in the prompt — photos, posters, logos, infographics — with about a third asking for exact text in quotes and many in languages other than English. The 9B teacher rewrote all of them on one A100 in 2 hours 37 minutes, about USD 6.50. After quality filtering, 1,840 examples were selected. Training the 0.8B and 2B students took 12 and 18 minutes on an A10G, USD 0.75 for both; the 0.8B also ships as an 812 MB GGUF file for CPU use. The project took about 11 hours and 24 jobs, at about USD 16. For Agate-Preview-002-4step, Logolabs' Agate Preview 002 is a 260M-parameter text-to-image model needing 50 steps with guidance, or 100 network passes per image; the agent distilled it to 4 passes. The first run cached 155,000 training images as latents, baked the guidance into the model and cut steps from 16 to 8 to 4 on A100s, with the 4-step student beating the teacher at the same 4 steps on GenEval and FID before ONNX export; that run took about 13 hours and USD 22. A second run generated 24,000 more image pairs at 16 steps and fine-tuned the student for about an hour, moving GenEval from 0.509 to 0.536 against the teacher's 0.563 at 50 steps. Total compute across both runs was about USD 37.
Additional coverage: Perplexity Brings Model Council to Computer — Turning Multi-Model AI Into Work-Ready Outputs
Hugging Face reported compute costs, project by project
The article publishes a cost table described as GPU and CPU job charges reported for each session. Citrus Doctor, on Qwen3.5-2B, is listed at USD 1.90. Huggy LoRA, on FLUX.2 klein base 4B, at USD 7.60. Pocket rewriter, on Qwen3.5-0.8B and 2B, at USD 16.05. Viewpoint Orbit LoRA, on Qwen-Image 2.1, at USD 16. Doodle-in LoRA, also on Qwen-Image 2.1, at USD 24.30. Agate 4-step across two runs, on Agate Preview 002, at USD 37. The stated total is about USD 103.
| Entity | Recent Focus | Geography | Source |
|---|---|---|---|
| Hugging Face | Hosts ML Intern in HuggingChat and the Hub where the resulting models, datasets and Spaces are published | Not stated in the source | Source |
| ML Intern | Plans work, requests budget, runs smoke tests, then trains, evaluates and publishes on Hugging Face hardware | Not stated in the source | Source |
| Qwen-Image 2.1 | Base model for the Viewpoint Orbit and Doodle-in LoRAs | Not stated in the source | Source |
| Qwen3.5-2B | Base model for the Citrus Doctor fine-tune, benchmarked before training | Not stated in the source | Source |
| FLUX.2 klein base 4B | Base model for the Huggy LoRA, also usable at 4 steps in distilled form | Not stated in the source | Source |
| Agate Preview 002 | Logolabs text-to-image model distilled from 50 steps to 4 | Not stated in the source | Source |
| Project-AgML | Hosts the three merged sources behind the citrus training dataset | Not stated in the source | Source |
Hugging Face Implementation Risks
The source describes a single author's experience and does not present controlled comparisons, so the reported accuracy and speed figures should be read as session results rather than broadly reproducible benchmarks. Project failure modes are acknowledged directly: the Viewpoint Orbit work took 48 jobs because runs failed on missing packages or wrong paths, and the Agate distillation required two runs. The budget gating depends on the agent requesting permission before paid jobs, which the author attributes to ML Intern starting at a zero dollar budget; that behaviour is described, not independently audited. Cost figures cover GPU and CPU job charges only and exclude the author's time and any costs outside the listed sessions.
Editorial independence disclosure: this article is based solely on the Hugging Face source cited below. Business 2.0 News has no commercial relationship with Hugging Face and did not receive input from the company or the authors on its content.
Source note: all facts above derive from the Hugging Face article at https://huggingface.co/blog/building-with-ml-intern . No additional sources were used.
What This Means for Practitioners
The practical reading for engineering teams is that the described constraint is prompt specification, not access to compute. Models such as Qwen3.5-2B and FLUX.2 klein base 4B are already on the Hub, and the reported spend per project stayed between about USD 1.90 and USD 37. What the account adds is a repeatable checklist: state the dataset, base model and training script; isolate already-verified facts so the agent does not re-derive them; require a zero-shot baseline before training; demand a smoke test with an explicit check before the full run; and set a hard budget cap. Teams evaluating agentic fine-tuning should pilot on a narrow task with a small cap, then inspect the published model card before scaling.
Analysis based on company announcements, investor disclosures, regulatory filings and publicly available market data as of publication.
About the Author
David Kim AI Author
AI & Quantum Computing Editor
David focuses on AI, quantum computing, automation, robotics, and AI applications in media. Expert in next-generation computing technologies.
David Kim is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →
Frequently Asked Questions
What is ML Intern?
ML Intern is an agent available in HuggingChat that, according to the Hugging Face account, plans work, asks the user for a budget before spending anything, runs a small test before the real job, then trains, evaluates and publishes on Hugging Face hardware.
How much did the models cost to build?
The source reports a total of about USD 103 in GPU and CPU job charges across the projects. Individual projects ranged from about USD 1.90 for Citrus Doctor to about USD 37 for the two-run Agate 4-step distillation. The author states these figures cover compute only, not time or underlying research.
What were the reported results of the first model?
The first project produced a 0.8B prompt rewriter that runs on a CPU, returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. Compute for the whole project, including the 9B model labelling 8,797 example requests, came to USD 16.
Which prompts does the author require in every project?
Two instructions are described as critical: a request for a zero-shot baseline before any training, and a smoke test with a check attached, such as 50 training steps plus verification that saved weights actually changed, before paying for the full run.
Are the reported accuracy figures independent benchmarks?
The source does not say so. The account describes a single author's sessions and does not present controlled comparisons, so the reported accuracy and speed figures should be read as session results rather than broadly reproducible benchmarks.