Technology

One
language for
computation.

Everything in KRAFT exists to answer one question: how does a piece of work safely leave the machine it was created on?

KRAFT IR

The
contract.

KRAFT IR is the universal description of computation — application-agnostic, hardware-agnostic, OS-agnostic.

An atom describes what is to be computed, where the inputs come from, which capabilities a node needs for it, how partial results belong together again and what happens on failure. No atom reads another atom's result during execution. That is the condition under which parallelism, reassignment and recovery are possible at all.

atom.json — IMAGE_TILE
{
  "atomType": "IMAGE_TILE",
  "operation": "tile_transform",

  "inputAssets": [{
    "type":     "IMAGE",
    "location": { "protocol": "http", "nodeId": "52c5c0f8" },
    "size":     67108864
  }],

  "capabilityRequirements": ["image_tile", "gpu"],
  "mergeStrategy":          "TILE_COMPOSITE",

  "recoveryPolicy": {
    "maxRetries":        2,
    "reassignOnFailure": true,
    "idempotent":        true
  }
}
Atom

IMAGE_TILE

Tile of an image. Merged via TILE_COMPOSITE.

Atom

FILE_RANGE

Byte range of a file. Hashing, scan, integrity.

Atom

DATA_BLOCK

Contiguous data range for CPU-bound parallel work.

Atom

VIDEO_CHUNK

Section of a video stream. Merged in frame order.

Atom

AUDIO_SAMPLE_BLOCK

Sample block. Stem rendering, offline effects, analysis.

Atom

TENSOR_BATCH

Batch for inference. Routed to Neural Engine or GPU.

Layers

Every
layer
has a
reason.

Take one away and computation can no longer be safely virtualized.

Two paths lead into the runtime: passive detection of running work and explicit submission through the gateway. Both meet in the Universal KRAFT Job — before it there is no IR, after it there is no way back.

L1
Work Detection
Observer, analyzer and detector watch running workloads. Result: work signals, and from them work atom candidates.
L2
Universal KRAFT Job
Candidates and gateway descriptors become the canonical, semantic job description: atom family, operation, inputs, split, merge and verification semantics.
L3
Adapter
Lowers jobs into strict KRAFT IR atoms. Adapters are software — they translate, they do not accelerate.
L4
IR Validator
Schema, bounds, capabilities, safety. No atom reaches the scheduler unchecked.
L5
Execution Planner
Decides how work is split and how large each piece becomes. Faster nodes get larger pieces.
L6
Scheduler
Picks the node per work unit — cost model, benchmark history, free resources, transport cost.
L7
Runtime Engine
Concurrency, retry, invariants, structured errors, health metrics. One entry point: runtime.submit(ir, chunks).
L8
Transport
Moves atoms and results — cost-aware, so the transfer is never more expensive than the gain.
L9
Merge & Verify
Partial results are reassembled and checked for integrity. Only then does a result exist.

No guessing.
No silent
failure.

IR Validator — layer 4

Live auto-offload

Nobody
presses
a
button.

KRAFT is not an export manager. KRAFT is a real-time runtime.

It does not wait for an export. It does not wait for a job. It does not become active on command and then go quiet again. A watcher sees new work the moment it appears and sends it down the path: Work Atom Candidate → Universal KRAFT Job → KRAFT IR → gateway.submit().

kraft — live offload
→ watcher    new file detected  frame_0871.tif  184 MB
→ candidate  atomizable       IMAGE_TILE
→ job        semantically confirmed
→ ir         valid              3 capabilities
→ planner    split              12 atoms · 64 MB
→ scheduler  52c5c0f8 ← 7   936ec7af ← 5
→ steal      936ec7af steals 2 from 52c5c0f8
→ merge      TILE_COMPOSITE     12/12
→ verify     ok                 sha256 matches
→ return     1.9 s              local: 5.4 s

Scheduling

Work
finds the
fastest
machine.

A static split is always wrong — at the latest when a node turns out slower than the plan assumed.

So atoms sit in a shared queue instead of fixed portions. A node that is done takes the next piece — across machine boundaries too. Whoever is fast works more. Whoever is slow holds nobody up. The result is identical, no matter who got which atom.

Cost model

Transfer must never cost more than the gain

Before an atom leaves the machine, the cost of transport is estimated. If it does not pay off, the work stays local.

Benchmarks

Measured, not estimated

Every run writes measurements back. The scheduler decides on history — not on data sheets.

Work stealing

A shared queue instead of fixed portions

Free nodes pull atoms from the shared queue. Every steal is logged and traceable in the report.

Zero-staging

Don't copy everything

Assets are not fully mirrored onto every node. A worker fetches exactly the range its atom needs over an HTTP range request.

Recovery

A failure is a retry

Atoms are idempotent. If a node fails, its atom is reassigned. The originating process notices nothing.

Verification

Reproducible or not at all

Computed across machines must deliver the same result as computed locally. If the checksum does not match, it is not a result.