Prose models can be trained
Training a prose model used to make it worse. The recommended training length was too long, and past about 2.1 passes over your examples a trained model lost to the untrained one on topics it was never taught. Prose now trains for rows × 0.7 rounds — 126 on a 180-row set. Extraction is unchanged at 400, because that number was measured on JSON extraction, where the model has an output format to learn.
How the training length was chosen
- One corpus trained at six lengths, with nothing else changed.
- A judge compared each trained model against the untrained one on 20 held-out rows about topics that appear nowhere in training, blind, reading every pair twice.
- The trained model wins up to 2.08 passes over the data and loses at 2.71 (n=20 at each length). The old default was 2.78. The new one is 0.69.
- A trained prose model is still not better than prompting the untrained one: asked to restate a taught fact in different words, the untrained model wins 12 to 8 (n=36). This release removes a default that made training harmful.
The test file you choose is the test file the run uses
- The wizard asked for a held-out file and then discarded it, carving 20% out of your training rows instead. On a 180-row dataset that is 144 training rows and 36 test rows; with a file chosen it is now 180 and 20, and your test set is used whole.
- The picker is reachable whichever way you added your data, and the job record keeps the path you chose with its checksum.
- A run refuses a test file your model was trained on, and refuses to overwrite a finished adapter.
- If you choose no file, BuildStudio holds back a slice of your own rows. On that dataset 0 of 36 test questions appear in training — but 36 of 36 test answers do, so the automatic split measures restating what was taught rather than handling something new.
Also
- The rounds control displayed 200 while holding 400, so the first drag halved your run. Fixed.
- On the command line, a note before training when a prose job asks for more than one pass over its rows.
- The Score screen counts how many held-out answers match the untrained model’s — 2 of 36 on our own model. It says the adapter changed the answers, not that it improved them.
Known limitations
- You can pick your training file as your test file. The app accepts it and the run refuses later; nothing bad gets through.
- An exported package runs in an Xcode app target. Built with
swift buildit compiles and then fails on a missing Metal library — copydefault.metallibbeside the executable, or build through Xcode. - Sidebar buttons expose no accessible name, 0 of 11. VoiceOver itself was not run.
- The refusal message names job-file keys, and one version of it leaves out the row count.
993b8ce. 6 of 8 release-gate checks passed, run by hand against the app inside this disk image; the one failure is the exported-package limitation above. buildstudio-0.6.3.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. The published file was downloaded and matched the local one by byte count and checksum. SHA-256: 2c9218ba803343878602be63458a56caff10321f567f9bb95dc7256881650107.