# Bansun on Replit

## Runtime

- Start with `pnpm --filter @workspace/bansun run dev`.
- Production build: `pnpm --filter @workspace/bansun run build`.
- The managed launcher owns loopback-only Ollama and the original Node web/API server. Do not start a second Ollama daemon.
- Default server models: `qwen2.5:1.5b` and `smollm2:1.7b`. Both original model sources are Apache-2.0 licensed. Qwen2.5 3B is research-only and is not a public server default.
- The launcher downloads missing models, verifies real one-token inference, then keeps them loaded while running. Read `/api/models` for `phase` and verified `warmed` models; installed files alone do not mean inference is ready.
- Two simultaneous server inference calls are allowed; extra requests receive a retryable busy response rather than an unbounded queue.

## Publishing

- Publishing is a separate user action; configuring a target does not publish the app.
- Target: Reserved VM. Start with at least 4 vCPU and 8 GB RAM for the two CPU models. Actual speed depends on the selected machine.
- Nix includes Ollama; the production build creates pinned browser-engine assets. Initial server model downloads need network access and several GB of disk.
- Keep credit enforcement and paid features disabled until credit state has durable storage. The original server credit keys/ledger and downloaded weights use local files, which are not guaranteed to survive republishing/redeployments. No database migration has been performed.
- Vault, recovery, journal and other private device records stay on the device. They are not uploaded as a hosting workaround. Export them before changing browser/origin or clearing storage.
- `$BANSUN` is not launched. There is no fabricated token address, payment, staking or purchase transaction.

## Browser/offline AI

- Nine choices: eight GPU models and a basic-English SmolLM2 135M CPU fallback.
- GPU models require a usable WebGPU adapter with shader-f16 and sufficient memory. Browsers without that hardware automatically select the CPU model.
- Loading a model is explicit. Public Hugging Face weights are revision-pinned; the browser CPU SDK/runtime are locally served, npm-version/integrity-pinned, with no CDN runtime fallback.
- After downloading, **Offline cache ready** means every recorded model file and required runtime asset is present. Loaded-in-memory alone is not an offline-reload guarantee.
- App updates preserve model/runtime caches. Browser eviction, storage clearing or changing devices can remove downloads.
- Offline chat cannot fetch live Ethereum, prices or wallet balances. Use online Agent tools for verified current data. Small models can make mistakes; source receipts and model confidence are not numerical validation.
- Model licenses differ. Review the linked original model terms; Qwen2.5 3B GPU weights are research-only.

## Readiness checks

- `pnpm --filter @workspace/bansun run build`
- `pnpm run typecheck`
- `pnpm audit --prod`
- `/api/health` must show the backend online. `/api/models` must verify warmed models, not just installed names.
- Verify browser inference again after an offline reload; a download badge or HTTP 200 for JavaScript is insufficient proof.
