Skip to content

Replace Espresso with a hardened ANE-LM backend - #86

Merged
IchenDEV merged 17 commits into
mainfrom
lava-topaz
Aug 31, 2026
Merged

Replace Espresso with a hardened ANE-LM backend#86
IchenDEV merged 17 commits into
mainfrom
lava-topaz

Conversation

@IchenDEV

@IchenDEV IchenDEV commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • replace the non-working Espresso dependency with the exact-pinned IchenDEV/ANE-LM@033472ec12ea796fc7ea4f8cefd7ed456f69900b SwiftPM runtime
  • add selectable local Qwen3 Hugging Face directory loading, native ANE generation, cancellation, reset, unload, and user-controlled MLX fallback
  • validate safetensors structure, required Qwen3 shapes and BF16 tensors, tokenizer byte-level coverage, contiguous IDs, merges, ChatML boundary tokens, and EOS mapping before saving or loading a model
  • release mapped native allocations deterministically and clean partial ANE state after failed loads
  • disclose private API, Mac App Store, compatibility, and delayed RSS-reclaim limitations in both localizations
  • use the macOS semantic under-page color for the shared Settings background

M5 evidence

  • official Qwen3-0.6B produced non-empty text through Utter on M5 Max / macOS 27 without MLX fallback
  • 20 requests passed across three complete load/generate/unload lifecycles, with one cancellation in the final lifecycle
  • within-lifecycle RSS growth was 80 KB, 32 KB, and 0 KB; every unload released more than 512 MiB from its peak
  • post-unload RSS remained nondeterministic, but final cross-lifecycle growth was 26,080 KB and below the enforced 128 MiB bound

This evidence is limited to Qwen3-0.6B on the tested M5 Max / macOS 27 host. Other model sizes, Apple Silicon devices, and OS versions remain unverified. The backend uses private Apple APIs and is not suitable for Mac App Store distribution; MLX remains the stable fallback.

Verification

  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer swift test — 592 XCTest tests passed, 10 environment-gated tests skipped, and the Swift Testing suite passed after merging current main
  • real ANE lifecycle test with UTTER_ANE_TEST_ITERATIONS=20 — passed in 84.367 seconds
  • bash scripts/sdlc-checks.sh — all 12 change bundles passed the strict shell gate
  • bash scripts/ci-basic-checks.sh — passed after conflict resolution and all approval-state updates
  • DEVELOPER_DIR=/Applications/Xcode-beta.app/Contents/Developer bash scripts/build-app.sh — release App and DMG built; app and mounted-DMG signatures and DMG checksum passed after merging current main
  • independent P1/P2 review — CLEAN

The two active intents, specs, plans, and verification artifacts were approved by the user on 2026-08-31. Merge and release remain separate human gates.

Evidence

  • docs/research/2026-08-31-apple-neural-engine-m5-compatibility.md
  • docs/sdlc/changes/2026-08-31-ane-lm-runtime/
  • docs/sdlc/changes/2026-08-31-settings-semantic-background/

@IchenDEV IchenDEV changed the title Add Espresso ANE inference backend Replace Espresso with a hardened ANE-LM backend Aug 30, 2026

@IchenDEV IchenDEV left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found three correctness/resource issues in the new ANE path and one defect in the environment-gated fallback test. The current smoke coverage can pass on partial Qwen reasoning output, long prompts silently exceed the native KV window, and a successful generation fallback leaves both local backends resident. These should be addressed before merge.

if modelName.lowercased().contains("qwen") {
return "<|im_start|>system\n\(system)<|im_end|>\n"
+ "<|im_start|>user\n\(user)<|im_end|>\n"
+ "<|im_start|>assistant\n"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Disable Qwen3 thinking for direct-output requests

Qwen3 treats this bare assistant\n suffix as its default thinking mode; its chat template only inserts the empty <think>\n\n</think>\n\n prefix when thinking is disabled. The formatting and edit-command paths expect direct text/JSON, use budgets as low as 256 tokens, and often pass temperature 0, which the pinned sampler implements as greedy decoding. The model can therefore spend the whole budget inside <think> before emitting the final answer, while the current real-model test still passes on any non-empty reasoning fragment. Please use the non-thinking suffix (or an equivalent hard /no_think switch) and assert final user-facing text/valid JSON in the real-model test.

ane_lm_generate(
model.runtime,
tokens.baseAddress,
tokens.count,

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject prompts that overflow the 2,048-token ANE cache

The pinned Qwen3 runtime has a fixed 2,048-entry KV cache and overwrites cache.start once it fills, while ane_lm_generate accepts every prompt length. Passing tokens.count here therefore makes prompts over 2,048 tokens silently forget their beginning before the first generated token. Screen context, memory, or longer transcripts can cross this even though the selected model advertises a much larger context. Guard the prompt length (ideally reserving output headroom) and fall back to MLX or truncate explicitly; add a >2,048-token regression test.

)
}
)
if result.usedMLX {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Unload ANE after a successful MLX fallback

If ANE loading succeeded but espressoLLM.generate throws, the actor still owns its native LoadedModel. This branch then returns a valid result after loading MLX, and the fallback policy only changes the saved backend to .mlx; it never runs the model lifecycle unload path. The process therefore keeps both full models resident until a later manual unload or exit, which can turn a recoverable ANE failure into memory pressure or an OOM. Unload espressoLLM before returning the successful MLX fallback (after preserving the failure message/outcome), and cover the generation-failure path with an isLoaded == false assertion.

)
if index == 0 {
baselineFootprint = currentMemoryFootprint()
options.localLLMBackend = .mlx

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve the fallback observation before switching iterations to MLX

With OPENTYPE_FALLBACK_ITERATIONS > 1, iteration 1 runs the .mlx branch, which calls clearEspressoOutcome(). Because the test consumes the tracker only after the loop, outcome is nil and XCTAssertEqual(outcome, .fallback) fails before the repeated-request memory check can validate anything. Consume/capture the fallback immediately after iteration 0, or keep a separate sawFallback flag.

@IchenDEV
IchenDEV merged commit bfa6216 into main Aug 31, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant