—— what’s new ——

change log.

Build small, task-specific models on your Mac, measure them honestly, and ship them as Swift packages.

current release  0.6.3

Prose models can be trained

Training a prose model used to make it worse. The recommended training length was too long, and past about 2.1 passes over your examples a trained model lost to the untrained one on topics it was never taught. Prose now trains for rows × 0.7 rounds — 126 on a 180-row set. Extraction is unchanged at 400, because that number was measured on JSON extraction, where the model has an output format to learn.

Download BuildStudio 0.6.3

How the training length was chosen

  • One corpus trained at six lengths, with nothing else changed.
  • A judge compared each trained model against the untrained one on 20 held-out rows about topics that appear nowhere in training, blind, reading every pair twice.
  • The trained model wins up to 2.08 passes over the data and loses at 2.71 (n=20 at each length). The old default was 2.78. The new one is 0.69.
  • A trained prose model is still not better than prompting the untrained one: asked to restate a taught fact in different words, the untrained model wins 12 to 8 (n=36). This release removes a default that made training harmful.

The test file you choose is the test file the run uses

  • The wizard asked for a held-out file and then discarded it, carving 20% out of your training rows instead. On a 180-row dataset that is 144 training rows and 36 test rows; with a file chosen it is now 180 and 20, and your test set is used whole.
  • The picker is reachable whichever way you added your data, and the job record keeps the path you chose with its checksum.
  • A run refuses a test file your model was trained on, and refuses to overwrite a finished adapter.
  • If you choose no file, BuildStudio holds back a slice of your own rows. On that dataset 0 of 36 test questions appear in training — but 36 of 36 test answers do, so the automatic split measures restating what was taught rather than handling something new.

Also

  • The rounds control displayed 200 while holding 400, so the first drag halved your run. Fixed.
  • On the command line, a note before training when a prose job asks for more than one pass over its rows.
  • The Score screen counts how many held-out answers match the untrained model’s — 2 of 36 on our own model. It says the adapter changed the answers, not that it improved them.

Known limitations

  • You can pick your training file as your test file. The app accepts it and the run refuses later; nothing bad gets through.
  • An exported package runs in an Xcode app target. Built with swift build it compiles and then fails on a missing Metal library — copy default.metallib beside the executable, or build through Xcode.
  • Sidebar buttons expose no accessible name, 0 of 11. VoiceOver itself was not run.
  • The refusal message names job-file keys, and one version of it leaves out the row count.
Verified: 258 of 285 Swift tests passed (27 skipped — they need a trained adapter and its held-out file on the machine) and 69 of 69 core tests passed, measured at 993b8ce. 6 of 8 release-gate checks passed, run by hand against the app inside this disk image; the one failure is the exported-package limitation above. buildstudio-0.6.3.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. The published file was downloaded and matched the local one by byte count and checksum. SHA-256: 2c9218ba803343878602be63458a56caff10321f567f9bb95dc7256881650107.

Train and export directly inside BuildStudio

Training no longer needs a separate Python environment. BuildStudio now runs the complete training path itself, preserves usable progress safely, and produces Swift packages that are ready to build and generate with.

Download BuildStudio 0.5.0

A simpler local training path

  • Training starts directly from the app without installing or selecting a Python environment.
  • Progress, cancellation, and checkpoints are managed by BuildStudio’s native training engine.
  • Checkpoints are written atomically so an interrupted run cannot leave a half-written adapter behind.
  • Only the answer portion of each example contributes to training, keeping the prompt context from becoming a target.

Packages that are ready to use

  • Exported packages locate their bundled model and adapter resources correctly at runtime.
  • Build verification no longer leaves a large hidden build directory inside the exported package.
  • Changes to generated package source and manifests correctly invalidate an older staged export.
  • The export screen reports completion only after the package has been saved and verified.
Verified: 72 Swift tests, 31 core tests, native training smoke coverage, and direct generation from an exported package. BuildStudio-0.5.0.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. SHA-256: 3cd9fd2e2d7044d785116ab4e17a6fcaaa631e38ea559dbe379d80319811e1d6.

Bring your own data—and use the model you trained

BuildStudio now accepts your existing paired dataset and reliably loads the complete adapter produced by training. You can start from a JSON or JSONL file instead of generating examples, then use the trained result in previews and exported Swift packages.

Download BuildStudio 0.4.1

Train with your own examples

  • The Data step now has a visible upload path for paired JSON and JSONL datasets.
  • Uploaded examples are checked locally against the answer contract before training.
  • A valid uploaded dataset goes directly to training; the independent quality review remains a gate only for examples generated by BuildStudio.
  • You can replace the uploaded file or switch back to generated examples without restarting the project.

Adapters load all the layers they trained

  • BuildStudio reads the saved adapter itself to find every trained projection instead of relying on a short hard-coded list.
  • Qwen attention and feed-forward adapters now attach correctly in native MLX inference.
  • The same loader ships in exported Swift packages, so app preview and packaged inference use the same trained weights.
  • Incomplete adapter pairs fail with a clear error before the model is changed.
Verified: 69 Swift tests, 31 core tests, and real Qwen and LFM adapter attachment checks. BuildStudio-0.4.1.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. SHA-256: d5ab168e…2497bb0.

Use the ChatGPT account you already have

BuildStudio no longer needs a separate assistant provider or account. Planning, example generation, and data checks now run through Codex using the ChatGPT sign-in already owned by the Codex CLI on your Mac.

Download BuildStudio 0.4.0

Simpler assistant setup

  • No OpenRouter key or BuildStudio assistant registration is required.
  • The Codex Account window shows sign-in status and can open browser login or sign out.
  • Each request gets a fresh, short-lived thread so previous reviews cannot bias the next one.
  • BuildStudio isolates its Codex configuration from unrelated MCP servers, plugins, and global instructions while securely reusing the existing login.
  • Versioned Guide and Critic prompt packs now ship inside the app.

Local work stays local

  • Training, scoring, preview, model storage, and Swift package export stay on the Mac.
  • Plans, generated examples, and second opinions send the relevant prompt and sample to OpenAI through the signed-in ChatGPT account.
  • Assistant usage follows that ChatGPT plan and requires the Codex CLI plus a network connection.
  • Assistant guidance is not yet on-device.
Verified: 64 Swift tests and a live signed-in Codex critique. BuildStudio-0.4.0.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. SHA-256: a6de2072…f65257d.

More models for every size of task

BuildStudio now gives you a broader set of Bonsai and LFM models. Pick a tiny model for fast, narrow work, a long-context LFM for larger inputs, or a larger Bonsai when the task needs more capacity.

Download BuildStudio 0.3.0

New choices

  • Ternary Bonsai now comes in 1.7B, 4B, and 8B sizes.
  • LFM2.5 choices include 230M, 350M, and 1.2B Instruct.
  • Each model explains what it is good at, where it is limited, and when it needs more memory.
  • Gemma is intentionally not part of this release; Bonsai and LFM are the supported families.

One model from training to shipping

  • The selected model carries through training, Try your model, held-out scoring, and Swift package export.
  • Adapters load directly without a conversion step.
  • Installed builds include a setup helper for local training.
  • A short two- or three-step smoke run is enough to verify the training path.

Known limitations

  • Larger Bonsai models need more unified memory and are not suitable for every Mac.
  • Bonsai 1-bit variants remain marked coming soon.
  • This release requires an Apple-silicon Mac running macOS 14 or later.
Verified: 64 Swift tests, 30 core tests, and smoke training on both Bonsai and LFM model families. BuildStudio-0.3.0.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. SHA-256: a0e2b401…1957f5.

From an idea to a reusable model

BuildStudio now designs the task with you before it generates data. The opening is a conversation that ends with an agreed task, a strict answer contract, and three examples from your domain. Completed models are now preserved in a library, so testing or packaging one never requires another training run.

A conversation, not a form

  • Luna asks one focused follow-up when the requested job is unclear.
  • The agreement includes a plain-language task, structured output fields, and three examples.
  • The third example covers missing information, an awkward input, or an out-of-scope case.
  • Users can refine the task, accept it, continue anyway, or start over without leaving the page.

Define the answer before creating data

  • The Data step is now Define the answer → Create examples → Review and approve.
  • Fields can come from the opening conversation, a separate Luna suggestion, an existing paired dataset, or a blank contract.
  • Names, types, descriptions, enum values, normalization, and evidence rules become one durable schema shared by generation, validation, training, scoring, and export.
  • The before-and-after test carries forward a generated example from the actual task instead of showing an unrelated built-in receipt.

Strict, reviewable generation

  • The assistant service now requests strict JSON Schema output with the exact approved fields, field types, and row count.
  • Large datasets are assembled in bounded 20-row batches; recent examples are supplied to later batches to reduce repetition.
  • Seed examples can be edited, removed, regenerated, or accepted despite a review finding.
  • Reviewer findings feed directly into regeneration instead of asking the user to manually recreate the same advice.
  • Creating the full dataset has a separate confirmation. It no longer begins from an incidental click.
  • Missing facts can stay empty rather than being replaced by a confident placeholder.

Models survive the wizard

  • A new Trained models library is available from the welcome screen and from every wizard step.
  • Each entry shows the base model, example count, training steps, adapter size, and date.
  • Any preserved model can be tested, packaged, or revealed in Finder immediately.
  • Future training runs use distinct durable directories instead of replacing the single studio run.
  • Relaunch recovery restores the original base, dataset, schema, and training settings from the resolved run record.

Stopping no longer throws away useful work

  • Training writes crash-safe checkpoints more frequently.
  • Stopping adopts the newest stable checkpoint when one exists and says when no checkpoint was ready.
  • A later export can reuse a matching interrupted checkpoint without secretly launching training again.

Safer Swift package export

  • BuildStudio asks where to save the package before packaging starts.
  • The private Application Support directory is now only a staging area; the verified package is copied to the chosen folder.
  • Existing folders are never silently overwritten.
  • Recovered models package with their original provenance, not defaults from a new wizard.
  • The static library bundled for export is guaranteed to match the release engine linked by the app.

Reliability and diagnostics

  • Unreadable assistant responses fail visibly and remain traceable in the assistant log.
  • Structural validation, reviewer approval, and final-dataset approval are separate gates.
  • Prompt pack sp-2026-08-04h adds the three-example planning contract and explicit output schema.
  • Assistant routes omit unsupported sampling parameters when strict structured output is required.

Known limitations

  • Luna’s planning, generation, and review need a network connection. Training, evaluation, inference, model storage, and package creation remain local.
  • Synthetic examples are a reviewed starting point, not a substitute for representative production data.
  • BuildStudio versions before 0.2.0 reused one run directory. Only the final intact adapter from that directory can be recovered.
  • This release requires an Apple-silicon Mac running macOS 14 or later.
Verified: 64 Swift tests, 29 core tests, strict assistant API tests, and a live three-example planning contract. BuildStudio-0.2.0.dmg is Developer ID signed, Apple notarized, ticket stapled, and Gatekeeper accepted. SHA-256: 1e1004d8…b90b7de7.

The complete local training loop

The first packaged BuildStudio release connected fine-tuning, held-out scoring, on-device inference, and Swift package export in one Mac app.

Included

  • A self-contained base model and native Apple-silicon training engine.
  • Paired JSONL validation, configurable LoRA training, live loss and memory reporting.
  • Base-versus-adapter scoring on held-out examples.
  • Developer ID signing, Apple notarization, and offline-capable Swift package export.
  • The first Luna second-opinion service for task fit, data quality, evaluation coverage, and ship decisions.