# File Upload SDK Framework-independent browser file/source lifecycle management. Sites declare contracts and render snapshots; this package owns transport, storage, identity checks, confirmations, deletion, reconciliation and workflow recovery. Public contract: https://file-upload-sdk.pages.dev/example. The Pages example is static documentation; product Workers still authorize requests and own server data. No API keys or session tokens enter local storage. ## Managed selections ```js import { createFileManager } from 'file-upload-sdk'; const manager = createFileManager(config); const unsubscribe = manager.subscribe(snapshot => render(snapshot)); await manager.restore(); await manager.addFiles(selectedFiles); await manager.addLinks([{ kind: 'web', url: 'https://example.org/article' }]); await manager.markPending(); // Persist explicit intent before sign-in/navigation. await manager.run(verifiedAccountId); // Also: remove(id, verifiedAccountId), retry(id, verifiedAccountId), // reconcile(verifiedAccountId), confirmPartial(), cancelPartial(), cancel(), dispose(). ``` `config` declares storage (`memory` or `indexeddb`), namespace/lock, admitted kinds, MIME/URL allowlists, file/total/count limits, request ID prefix, optional image normalization, public auth-controller subscription, messages and HTTP operations. No site-provided callbacks decide mutation outcomes or recovery. Each operation declares a same-origin path, method, bounded deadline and response size. Request bodies support JSON, multipart or raw bytes. `{$ref:'contextId'}` selects an operation parameter; `{$template:'Item {draftId}'}` formats text. Response descriptors select `result`, validate consumed fields with `schema`, check envelope `success` and identity `matches`, then optionally `projection` renames fields. Unknown fields are not consumed. Missing required values fail. Multipart operations also accept an ordered array of `{name,value,filename?}` parts, preserving repeated field names and original file bytes. A native `FormData` can be passed through a `{$ref:...}` descriptor. `{$json:descriptor}` serializes a resolved value for a multipart JSON field; ordinary `$ref` values are never implicitly JSON-encoded. Configure `successStatus:202` for a job admission rather than treating acceptance as completed processing. `headerMatches` checks configured response headers before accepting either JSON or binary responses. Schemas can bound string lengths and array counts, and enforce finite numeric values, integer byte counts and numeric ranges. These checks complement response byte limits and exact request correlation. Operations: create/verify a context, complete source lists, imports by kind, download, remove, prepare-one, preparation report, and generation. Binary reads return `{blob,headers}`. Generation projects its proof into `{fingerprint}`; preparation reports contain `{extractions:[{sourceId,fingerprint}]}`. The SDK checks exact source IDs and SHA-256 of sorted unique preparation fingerprints before clearing a selection. It validates again after generation. The complete deployed-product configuration is maintained in ../capyrecap-website/src/lib/file-manager-config.ts. Keep new products declarative; never copy SDK state machines into their components. ## Ownership and recovery Sources retain local IDs, original files/URLs, optional normalized upload files, SHA-256/URL identities, server receipts and `pending/importing/imported/failed/uncertain` states. Drafts retain revision, owner, context and workflow progress. Every mutation rereads under its configured Web Lock; unavailable IndexedDB/Web Locks fail visibly. Memory-only selections need no browser persistence. The SDK queues file batches, reports individual rejection reasons and retains a batch in memory when persistence fails. Resume it explicitly with `retry()`. Partial processing requires `confirmPartial()`, no uncertain sources, and at least one confirmed receipt. A changed selection invalidates confirmation. Transport cancellation, timeout, HTTP failure or malformed success never proves absence of a server write. Only an explicitly configured, contract-proven `rejectedBeforeWrite` tuple does. Otherwise reconcile through complete owned lists and downloaded byte hashes/exact URLs; never match filenames or request IDs. A failed or ambiguous read preserves uncertainty. No automatic mutation replay. Imported file removal verifies ownership, then deletes the exact server ID. A lost reply is resolved by a complete server list; failed verification keeps the receipt. Local removal occurs after confirmed deletion/absence only. A local persistence failure can be retried without uploading or deleting an unrelated file. `restore()` never starts uploads or generation. `markPending()` records the user's explicit intent. `run()` checkpoints context creation before sending; unknown creation cannot silently create another context. Confirmed imports survive failed preparation/generation. Account changes cancel protected operations when an auth controller subscription is configured; ownership remains enforced server-side. ## Storage migration Schema version 3 is current. A configured `storage.migration` declares accepted prior versions and the prior context-field name. Migration preserves original files, identifiers, ownership, receipts and revisions; old unconfirmed imports become uncertain. Unsupported/corrupt schemas remain stored and return errors. CapyRecap declares its existing database and lock names and migration from unversioned/version-2 drafts using `contextField:'projectId'`. ## Stateless and transient requests `upload(options)` remains the bounded multipart/raw XHR primitive for consumers requiring byte progress only. It never interprets 100% sent as persisted success. `createContractClient(prefix, getSignal)` handles configured JSON/blob requests, correlation, schemas, deadlines and cleanup. It is the upload-only/transient mode: no local draft, deletion endpoint or persistent server resource is required. `createSourceClient(config)` shares admission and image normalization with managed selections. Its `prepare(context, progress, signal)` sequences per-source analysis and checks the final report. `generate(context, progress, signal)` also verifies the generation fingerprint and rereads the complete source set. These stateless actions are suitable for existing server contexts and one-shot covers; they do not persist local recovery receipts or automatically resend mutations. Durable selections use `createFileManager`. URL admission can declare permitted hosts and identifier selectors (`query`, `pathPrefix`, `segment`, `hosts`) with an `identityPattern`; products never implement their own video-link parser. `file-upload-sdk/preview` exports `createPreviewLoader(config)`. It owns preview scheduling, PDF/image rasterization, byte transport and resource cleanup. PDF.js is an optional peer; importing the core does not load it. Local originals remain untouched; preview endpoints must not persist the preview as an imported source. `FileManagerError` carries code, status, operation and requestId. Source snapshots also expose safe display errors, imported/excluded sources, removable/retryable IDs, partial confirmation, queued files and truthful phase/byte progress. `UploadError` retains its transport-only format. No tokens, file contents or transcripts appear in SDK error diagnostics. ## Asynchronous processing `createProcessingClient(config)` handles multipart admission and asynchronous processing without requiring preparation or generation fingerprints. Its configuration declares create/read/remove/download operations, input fields, file limits, status values, output identities and the server's idempotency contract. The SDK preserves original bytes; it does not render or modify a Fabric canvas. Fabric exports or selected original files are ordinary inputs. ```js const jobs = createProcessingClient(config); const receipt = await jobs.run({ fields: {capability_id: modelId, catalog_type: category, params}, files: [{field: 'input', file: originalFile}], }, {ownerId: accountId, signal, onProgress, onStatus}); const output = await jobs.download(receipt.outputs[0].url, {ownerId: accountId}); // restore(ownerId), retry(checkpointId, options), remove(checkpointId, options) // subscribe(listener), cancel(), dispose() ``` Checkpoints use a separate version-1 store in configured memory or IndexedDB. IndexedDB requires a named database and Web Lock namespace. Each input fingerprint has its own lock: distinct files may run concurrently; the same pending request cannot be submitted by two tabs. Checkpoints retain the owner, stable UUID, declared fields, original files and filenames, state and receipts. Authentication credentials are supplied only through ephemeral options and never saved. An auth-controller subscription cancels protected work on account change. The SDK saves before sending. A lost admission response triggers an owned read for that UUID; a matching receipt resumes polling without another POST. `run()` reuses unfinished identical inputs, including after reload. `restore()` reads local checkpoints without network requests or automatic processing. Explicit `retry()` may replay an unacknowledged admission only when a correlated, validated error matches the configured absence status and code, and the server guarantees atomic `same-id-and-body` replay. Previously accepted or expired jobs are never recreated by this mechanism. A failed read preserves uncertainty. `rejectedBeforeWrite` declares exact status/code pairs whose backend contract guarantees rejection before admission. Only a validated, correlated error may mark a checkpoint `rejected`; a generic HTTP status never suffices. Explicit retry retains original bytes and UUID. Other failure responses remain uncertain. Credential descriptors use `credentials.` and cannot replace checkpoint identifiers, multipart parts or saved input fields. Success requires a terminal receipt with exact job identity, unique output slots, job-bound output paths and available outputs. Downloads reread the owned terminal receipt and check the selected output path and returned MIME type. Reads and binary bodies have independent bounds within the total processing deadline. Failures and cancellation retain checkpoints. Deletion clears the local record only after a configured deletion proof or a verified server absence. Server completion and browser delivery are separate. Snapshots expose `delivered` and `deliveredSlots`; a completed job whose outputs were not delivered is reused by the next identical `run()`, including after reload. Successful, nonempty, MIME-verified downloads atomically checkpoint each slot. Only delivery of every reserved output permits a fresh identical run. JSON-only jobs complete delivery with their validated terminal receipt. Cancellation or account change prevents a late result from being published. Presentation callbacks receive independent snapshots and cannot modify recovery proofs. Related product: ../editgator-website/src/features/processing-upload-config.js owns routes, limits and proofs; processing-job-api.js owns account activation, presentation errors and conversion to existing tool result shapes. The private ../workers/editgator/editgator-api/ owns authorization, credit reservation, idempotent admission, queued execution and artifact storage. Browser-only Fabric, image, PDF, audio and video capabilities remain local and create no remote jobs. ## Content-addressed assets `createAssetClient(config)` manages original files that become reusable stored references, without starting a processing job. `upload(file, fields, options)` persists original bytes before admission and returns only after reading back the stored bytes and checking their size and SHA-256. Repeating the same input reuses its verified reference. Canvas import remains local; the host invokes upload only for a save or an explicit remote action. Declare `contentTypes` as MIME-to-extension mappings, `maxFileBytes`, a total `timeoutMs`, bounded `inputSchema`, independent IndexedDB database/Web Lock, authentication subscription, and `operations.create/lookup/download/remove`. Every operation uses the same response/schema/identity descriptors as the other SDK clients. The operation context contains `ownerId`, `id`, original `file`, `hash`, `bytes`, `contentType`, `extension`, `fields`, and runtime `credentials`. The declarative `receipt` is resolved from those verified values after the owned lookup and binary read have succeeded. Lookup results are available as `stored.` for the binary-read request and receipt, so an API may issue opaque references without the browser constructing storage addresses. Saved receipt values are available as `receipt.` for configured deletion. ```js import {createAssetClient} from 'file-upload-sdk'; const assets = createAssetClient(assetConfig); const receipt = await assets.upload(originalFile, {canvasId}, { ownerId: verifiedAccountId, onProgress: renderProgress, }); // restore(ownerId) reads local state without making requests. // retry(checkpointId, options) explicitly resumes uncertain uploads. // remove(checkpointId, options) confirms server deletion/detachment first. // cancel()/dispose() abort active operations and fence account changes. ``` A lost upload response triggers an owned lookup and exact binary comparison, never a second automatic upload. Unverified writes remain recoverable after reload. `idempotency.replay:'same-content-address'` permits explicit retry only after a correlated canonical absence with the configured code/status. Declare this guarantee only when the backend admits identical original bytes under one immutable owner-bound address. Give asset and processing clients separate storage databases. Authentication credentials are never saved. ## Publication and validation commands Maintain plain JavaScript source and TypeScript declarations; no local build. Storage suites use the development-only `fake-indexeddb` fixture. Run packaged `tests/*.test.js` remotely through EditGator and CapyRecap Full gates, alongside its configuration and product integration tests. Browser acceptance also checks real IndexedDB, reloads, multiple tabs and interrupted network operations. Pack a new immutable version with `npm pack --ignore-scripts`, update the consumer's archive/dependency/lockfile, and publish each owning repository with `github-publish-main`. Existing sites pinned to 1.0.0 are unaffected. Publication on canonical `main` triggers the remote Pages build. Verify its exact commit, terminal deployment result and GET /example. Server APIs are unchanged by this migration.