# Osolix · Provisional Patent Application — Invention Disclosure
## Hybrid Online/Offline AI Resolver for Multi-Tenant Enterprise Software

> **Owner:** Senior Copyright & Intellectual-Property Officer
> **Sponsor:** Founder + Head of AI Development
> **OKR target:** 2026-Q3 KR-2.3 — file 1 provisional patent application covering the Hybrid Online/Offline AI architecture
> **Status:** Draft v1 — review-ready for outside counsel
> **Filing target:** US provisional + Madrid Protocol (UAE / GCC priority) within 60 days

---

> ⚠️ **Confidentiality:** This document is privileged invention-disclosure
> material under the Osolix LLC employee + contractor IP-assignment
> agreements. Distribution outside the Osolix legal team and the named
> outside patent counsel requires written consent from the Founder.
> A clean version (without trade-secret coefficients) is published as
> the public-facing marketing claim.

---

## 1 · Title of the invention

**System and method for parallel-tier AI inference with confidence-based answer selection in multi-tenant enterprise software, including offline-first execution within a tenant boundary.**

Working title (short): *"Hybrid AI Resolver."*

---

## 2 · Inventors

- **Anas Abu Ehmade** — Founder · primary inventor of the multi-tenant
  application of the parallel-tier resolver pattern + the offline-first
  Sovereign-tenant variant
- **Head of AI Development** — co-inventor of the JSON-salvaging parser
  for small open-weights model output + the compact-prompt fallback for
  sub-7B parameter models
- *(Outside patent counsel will refine the inventor list during the
  formal filing per US 35 U.S.C. §116.)*

---

## 3 · Field of the invention

The invention lies in the field of **multi-tenant enterprise software-as-a-service (SaaS) systems that integrate Artificial Intelligence (AI) inference into business workflows**, specifically:

- Fixed Asset Management (FAM) systems
- Enterprise Resource Planning (ERP) systems
- Document-extraction systems
- Multi-tenant SaaS platforms requiring per-tenant AI configuration

More narrowly, the invention concerns **AI inference architectures that run multiple model tiers in parallel on a single user request and select the response with the highest confidence score**, with provisions for tenant-specific data-sovereignty constraints.

---

## 4 · Background of the invention

### 4.1 The state of the art

Existing multi-tenant SaaS platforms incorporate AI inference under one of two predominant architectures:

**Architecture A — Cloud-AI-only** (exemplified by Mojodat · Asset Panda · MaintainX). Every AI-driven request routes to a centrally-hosted, vendor-controlled large language model (LLM) — typically OpenAI, Anthropic, or Google. The platform charges per-call, per-token, or via a flat AI-add-on subscription. Tenant data egresses to the LLM provider for every request.

**Architecture B — Local / no-AI** (exemplified by Snipe-IT · legacy on-premise FAM). The platform either omits AI inference entirely or relies on rule-based heuristics and historical-data lookups. No cloud API calls are made; tenant data does not leave the tenant boundary.

### 4.2 The problem

Architecture A imposes three structural costs on the customer:

1. **Per-call AI cost** scales linearly with asset / transaction volume. At scale (e.g. 50,000+ assets) the AI-add-on becomes a measurable line item.
2. **Tenant data egress** to the LLM provider is contractually mitigated but technically present, creating issues for regulated tenants (government · defence · banking · healthcare).
3. **Vendor lock-in** to the LLM provider's pricing, model deprecation schedule, and API availability.

Architecture B avoids those three costs but loses access to the materially higher accuracy of state-of-the-art cloud LLMs on tasks like invoice extraction from low-resolution scanned PDFs, semantic asset classification, and predictive maintenance.

The customer is forced to choose. The choice is binary, ex-ante, and difficult to revisit without re-platforming.

### 4.3 The unmet need

There is no known prior art that combines:

- **Per-request parallel invocation** of both a cloud-hosted LLM (online tier) AND a tenant-hosted LLM (offline tier)
- **Confidence-score-based answer selection** between the two tiers, with the selection rationale recorded in the response
- **Per-tenant configurable tie-break preference** (online-favoured for accuracy-first tenants vs offline-favoured for sovereign-first tenants)
- **Graceful degradation** to whichever tier responds when the other fails
- **Audit-traceable logging** of every parallel invocation with both tiers' confidence scores per call
- **Within a multi-tenant SaaS platform** where the tenant administrator chooses the AI mode (online · offline · hybrid · sovereign) per workflow and per request type

The closest prior art (research preprints on "model cascading" and "ensemble inference") concerns single-tenant, research-grade architectures that select between models by request difficulty, not by confidence-tier-comparison after parallel invocation. None addresses the multi-tenant per-request selection problem with the configurable tie-break and audit trail described herein.

---

## 5 · Summary of the invention

The invention is a **resolver pattern** implemented in a multi-tenant SaaS platform that:

1. Receives a user request requiring AI-driven analysis (e.g. invoice extraction, asset classification, predictive maintenance scoring).
2. **Reads the calling tenant's AI mode configuration**, including:
   - Online · Offline · Hybrid · Sovereign mode flag
   - Tie-break preference (`PreferOfflineOnTie`)
   - Per-agent enable / disable flag
3. **In Hybrid or Sovereign mode**, dispatches the request **in parallel** to:
   - An online tier (cloud-hosted LLM, e.g. Anthropic Claude)
   - An offline tier (tenant-hosted LLM via OpenAI-compatible chat-completions API, e.g. Ollama, vLLM, llama.cpp, LM Studio)
4. Awaits both tier responses (typically returned in 25-90 seconds total wall-clock, dominated by the slower tier).
5. **Selects the response with the higher confidence score**, with tie-breaking governed by the tenant's preference flag.
6. **Records the decision in an audit log** including:
   - Both tiers' confidence scores
   - The selected tier (the "source")
   - The reasoning string explaining the selection
   - The tenant ID, the user, the IP, the user-agent
7. **Returns the response with the source badge stamped on every field** so downstream consumers (UI, downstream services, the user) see which tier produced each value.
8. **Falls back gracefully**: if either tier fails (network error, API rate limit, parse failure), the surviving tier wins automatically without re-prompting.

The novelty rests not in any single component (parallel invocation, confidence scoring, audit logging, multi-tenant configuration are all individually known) but in the **combination** specifically tuned to the multi-tenant SaaS use case with per-tenant data-sovereignty constraints.

---

## 6 · Detailed description

### 6.1 Architectural overview (referring to Figure 1 — described textually)

```
                ┌──────────────────────────────────────────────────────┐
                │  Tenant Configuration                                │
                │  ─ AiMode: Online | Offline | Hybrid | Sovereign     │
                │  ─ PreferOfflineOnTie: bool                          │
                │  ─ LocalVisionEndpoint, LocalVisionModel, ApiKey     │
                │  ─ Per-agent enable/disable flags                    │
                └──────────────────────────────────────────────────────┘
                                         │
                                         ▼
                              ┌──────────────────────┐
   user request ─────────────►│  Resolver Service    │
   (e.g. invoice              │  (per-request scope) │
   extraction)                └──────────────────────┘
                                         │
                       ┌─────────────────┴─────────────────┐
                       │                                   │
                       ▼                                   ▼
            ┌────────────────────┐              ┌────────────────────┐
            │  Online Tier       │              │  Offline Tier      │
            │  ─ Anthropic API   │              │  ─ Tenant-hosted   │
            │  ─ Claude Sonnet   │              │    OpenAI-compat   │
            │  ─ Vision + text   │              │    chat completion │
            │  ─ Returns confi-  │              │  ─ JSON salvager   │
            │    dence + payload │              │  ─ Compact prompt  │
            └────────────────────┘              └────────────────────┘
                       │                                   │
                       └─────────────┬─────────────────────┘
                                     ▼
                          ┌────────────────────┐
                          │  Selection Logic   │
                          │  ─ pick higher cf  │
                          │  ─ tie-break per   │
                          │    tenant flag     │
                          │  ─ surviving tier  │
                          │    wins on failure │
                          └────────────────────┘
                                     │
                       ┌─────────────┴─────────────┐
                       ▼                           ▼
            ┌────────────────────┐      ┌────────────────────┐
            │  Audit Log         │      │  User Response     │
            │  ─ Both tiers' cf  │      │  ─ Winning payload │
            │  ─ Selected source │      │  ─ Source badge    │
            │  ─ Reasoning       │      │    per field       │
            │  ─ Tenant + user   │      │  ─ Confidence      │
            └────────────────────┘      └────────────────────┘
```

### 6.2 The resolver service

The resolver service (referred to in the implementation as `HybridInvoiceExtractorCopilot` and analogous classes for asset classification, maintenance, retirement, and chat) executes the following pseudocode:

```
function Extract(tenantId, payload):
    config = LoadTenantAiConfig(tenantId)

    if config.AiMode == Offline:
        return OfflineCopilot.Extract(tenantId, payload)
    if config.AiMode == Online:
        return OnlineCopilot.Extract(tenantId, payload)

    // Hybrid + Sovereign: parallel invocation
    onlineTask  = SafeAsync(() => OnlineCopilot.Extract(tenantId, payload))
    offlineTask = SafeAsync(() => OfflineCopilot.Extract(tenantId, payload))
    preferOfflineOnTie = config.PreferOfflineOnTie

    await Task.WhenAll(onlineTask, offlineTask)

    return PickBest(onlineTask.Result, offlineTask.Result, preferOfflineOnTie)
```

The `SafeAsync` wrapper ensures a thrown exception in either tier produces a `null` result rather than failing the entire request, enabling the graceful-degradation behaviour described above.

### 6.3 The selection logic

```
function PickBest(onlineRes, offlineRes, preferOfflineOnTie):
    if onlineRes is null and offlineRes is null:
        return ConstructFallbackResponse(reason: "both tiers failed")

    if onlineRes is null:
        return WithHybridSource(offlineRes, "Offline won by default")
    if offlineRes is null:
        return WithHybridSource(onlineRes, "Online won by default")

    if offlineRes.Confidence > onlineRes.Confidence:
        winner = offlineRes
    elif onlineRes.Confidence > offlineRes.Confidence:
        winner = onlineRes
    else:
        // exact tie
        winner = preferOfflineOnTie ? offlineRes : onlineRes

    return WithHybridSource(winner, EnrichReasoning(winner, onlineRes, offlineRes))
```

`WithHybridSource` enriches the response's reasoning string with the comparative confidence ("Hybrid: Online tier won — confidence 0.91 vs Offline 0.84") so the consumer can audit the selection rationale.

### 6.4 The offline tier — supporting innovations

The offline tier is invoked over the **OpenAI-compatible chat-completions API** which is the de-facto interoperability standard for self-hosted LLM runtimes (Ollama · vLLM · LM Studio · llama.cpp server). The same prompt + image payload that the online tier sends to Anthropic is sent (with adapter translation for the system-prompt convention) to the tenant-hosted runtime.

Two supporting innovations make the offline tier viable on the small CPU-friendly models that customers can practically host on commodity hardware (1.7B - 7B parameter range):

#### 6.4.1 JSON salvager

Small open-weights vision models (moondream 1.7B, llava-phi3 3.8B, qwen2-vl:3b) routinely emit JSON output wrapped in conversational prose, e.g.:

> "Here is the structured invoice data you requested: { ...JSON... } Let me know if you need anything else!"

The salvager extracts the JSON object from this prose by:
1. Stripping markdown fence markers (` ``` ` and ` ```json `)
2. Locating the first `{` character
3. Walking forward with a depth-counter that ignores braces inside string literals
4. Returning the substring from the first `{` to the matching closing `}`

This recovers structured output reliably from sub-7B parameter models that the online tier (Anthropic Claude) handles trivially because of its much higher prompt-following fidelity. The salvager is a small but specific innovation that materially expands the set of offline models the platform supports.

#### 6.4.2 Compact-prompt fallback

When the configured offline model is small (matched by name regex against the model identifier — `moondream`, `llava-phi`, `phi-3.5-vision`, `:3b`, `:1.7b`, etc.), the resolver swaps the full multi-rule extraction prompt (≈50 lines) for a tight single-imperative-sentence prompt (≈3 lines) plus a single output-shape example. Small models follow the compact prompt with materially higher reliability. The compact prompt is generated once per release and stored as a constant; the model-size-based selection runs per-request at zero cost.

### 6.5 Audit trail

Every parallel invocation writes one row to the platform's immutable audit log table with:

- `TenantId` · `UserId` · `UserEmail`
- `Timestamp` (UTC)
- `EntityType` = `"Auth"` or `"AI"` depending on workflow
- `Action` = `"AI Extracted (SSO · Hybrid)"` or analogous
- `NewValues` JSON containing both tiers' confidence scores, the winning tier, and the reasoning string
- `IpAddress` · `UserAgent`

The audit row is the canonical evidence for compliance audits (SOX 404, ISO 27001) showing how each AI-driven decision was reached. The dual-tier comparison in the audit row is the specific element that no prior-art audit log captures.

### 6.6 Tenant configuration

The tenant administrator selects the AI mode from a 4-option dropdown in the platform's admin UI:

| Mode | Online tier | Offline tier | Tie-break | Use case |
|---|---|---|---|---|
| Online | Always invoked | Never | n/a | Cost-tolerant, accuracy-first |
| Offline | Never | Always | n/a | Sovereign / air-gapped tenants |
| Hybrid | Both | Both | Online wins on tie | Best-of-both for non-regulated tenants |
| Sovereign | Both | Both | Offline wins on tie | Regulated tenants with optional cloud assist |

The mode is stored on the tenant record as an enum + boolean tie-break flag. Per-agent enable/disable flags allow the tenant to disable AI on specific workflows (e.g. retirement-of-high-value-asset) while keeping it enabled elsewhere.

---

## 7 · Claims (broad → narrow)

> **Note:** The following claims are review-ready for outside counsel. The
> exact dependent-claim structure will be optimised by counsel; the
> substantive scope is captured below.

### Claim 1 (broad — system claim)

A multi-tenant software-as-a-service system comprising:

- a tenant configuration store storing, for each of a plurality of tenants, an AI-mode flag selectable among at least online, offline, and hybrid modes;
- a request handler configured to receive AI-inference requests and read the calling tenant's AI-mode flag;
- a resolver service configured, when the AI-mode flag indicates hybrid mode, to dispatch the request in parallel to:
  - an online tier comprising a cloud-hosted large language model accessed via remote API; and
  - an offline tier comprising a tenant-hosted large language model accessed via a local network endpoint;
- a selection logic configured to receive the responses from both tiers and select the response having the higher confidence score, with tie-breaking governed by a per-tenant tie-break preference flag;
- an audit log configured to record, for each request, the confidence scores from both tiers, the selected tier, and a reasoning string explaining the selection.

### Claim 2 (dependent — Sovereign mode)

The system of claim 1, wherein the AI-mode flag is selectable among online, offline, hybrid, and sovereign modes, and wherein the sovereign mode comprises hybrid-mode parallel invocation with the per-tenant tie-break preference flag set such that the offline tier wins on equal-confidence ties.

### Claim 3 (dependent — graceful degradation)

The system of claim 1, wherein the resolver service is further configured such that, if either the online tier or the offline tier fails to respond or returns a parse error, the surviving tier's response is selected automatically without re-prompting.

### Claim 4 (dependent — JSON salvaging)

The system of claim 1, wherein the offline tier's response handling further comprises a JSON salvager configured to extract a structured object from a response containing prose-wrapped JSON, by:

- stripping markdown fence markers from the response;
- locating the first opening brace character;
- walking forward through the response with a depth counter that ignores brace characters appearing within string literals; and
- returning the substring from the first opening brace to the matching closing brace.

### Claim 5 (dependent — compact-prompt fallback)

The system of claim 1, wherein the resolver service is further configured to, before dispatching a request to the offline tier, examine an identifier of the offline-tier model and, if the identifier matches a pre-defined set of identifiers indicating a small-parameter-count model, substitute a compact prompt for the standard extraction prompt prior to dispatch.

### Claim 6 (independent — method claim)

A method, performed by a multi-tenant software-as-a-service system, of generating an AI-inference response, the method comprising:

- receiving an AI-inference request from a user associated with a tenant;
- reading an AI-mode flag associated with the tenant;
- if the AI-mode flag indicates hybrid or sovereign mode, dispatching the request in parallel to an online tier and an offline tier;
- receiving responses from both tiers, each response including a confidence score;
- selecting one of the responses based on the confidence scores and a per-tenant tie-break preference flag;
- recording, in an immutable audit log, both confidence scores, the selected tier, and a reasoning string;
- returning the selected response, with a source-of-answer indicator stamped on each field of the response.

### Claim 7 (dependent — audit log shape)

The method of claim 6, wherein the audit log entry further comprises the tenant identifier, the user identifier, the request IP address, and the user-agent string, and wherein the audit log table is configured such that records cannot be modified or deleted after creation.

### Claims 8-12 (additional dependent claims to be added by outside counsel)

Including but not limited to:
- Claim covering the OpenAI-compatible chat-completions API as the offline-tier transport
- Claim covering the per-agent enable/disable flag in tenant configuration
- Claim covering the use of vision-capable LLMs in both tiers for document extraction
- Claim covering the multi-tenant query-filter mechanism that prevents cross-tenant data leakage between AI requests
- Claim covering the cost-transparency banner shown to administrators in the admin UI

---

## 8 · Drawings list

The patent application will include the following drawings (figures):

- **Figure 1** — Architectural overview (referenced as text in §6.1)
- **Figure 2** — Resolver service flowchart
- **Figure 3** — Selection-logic decision tree including tie-break paths
- **Figure 4** — Audit log entry schema
- **Figure 5** — Tenant configuration UI mock-up (the AI-mode selector)
- **Figure 6** — Offline-tier supporting components (JSON salvager + compact-prompt fallback)

Outside patent counsel will produce the formal drawings to USPTO standards.

---

## 9 · Prior art search notes

The CISA Manager + outside counsel will conduct a formal prior-art search before filing. Preliminary review surfaced:

- **OpenAI / Anthropic API documentation** — describes single-tier inference, no parallel-tier resolver
- **Microsoft Azure OpenAI Service** — single-tier inference within an Azure tenant; no parallel cloud + on-prem
- **AWS Bedrock** — multi-model registry but routes each request to a single model; no parallel invocation with confidence-tier selection
- **Mojodat / Asset Panda / Maximo / Oracle FA** — published material indicates single-tier (cloud-only) AI integration; no offline parallel-tier
- **Academic papers on "model cascading"** — concern single-tenant research architectures that route by request difficulty rather than confidence-tier selection after parallel invocation
- **Academic papers on "ensemble inference"** — concern combining multiple models' outputs rather than selecting between them

Outside counsel will produce a formal Information Disclosure Statement (IDS) for the USPTO listing all relevant prior art.

---

## 10 · Filing strategy

| Filing | Jurisdiction | Target date | Rationale |
|---|---|---|---|
| **US provisional** | USPTO | 2026-Q3 mid-quarter | Establishes US priority date with low filing cost (~$300 fees + $5K-15K outside counsel for the full disclosure) |
| **PCT application** | WIPO | 2027-Q3 (12 months after US provisional) | Extends international priority window to 30 months for national-stage filings |
| **National-stage filings** | UAE Ministry of Economy + EPO + India + KSA + UK + Japan + Singapore | 2028-Q3 — 2029-Q3 (within 30 months of PCT) | Covers the markets where Osolix targets enterprise FAM customers |
| **Trademark filings (parallel)** | UAE Ministry of Economy + Madrid Protocol | 2026-Q3 | Osolix™, the gold-O monogram, "Sovereign AI for Fixed Assets™" |

### 10.1 Total filing budget (3-year)

| Year | Activity | Estimated cost (USD) |
|---|---|---|
| Year 1 (2026) | US provisional + outside counsel disclosure drafting | $20,000 — $25,000 |
| Year 2 (2027) | PCT filing + international search report | $15,000 — $20,000 |
| Year 3 (2028-2029) | National-stage filings in 5-7 jurisdictions | $80,000 — $150,000 |
| **Total** | **3-year IP investment** | **$115,000 — $195,000** |

### 10.2 Strategic rationale

The provisional patent serves three purposes:

1. **Defensive moat** — formal documentation that the Hybrid AI Resolver predates any future competitor's similar architecture
2. **Sales artefact** — "patent-pending" claim on the marketing site is a measurable enterprise-evaluation signal
3. **Acquisition signal** — clean IP estate increases enterprise value at any future M&A or strategic-investor event

---

## 11 · Inventor IP-assignment confirmation

By co-signing this disclosure document, each named inventor confirms:

- The invention was conceived in the course of their engagement with Osolix LLC under the standard employment / contractor IP-assignment agreement
- No prior or concurrent assignment of this invention exists to any other party
- The inventor agrees to cooperate with outside counsel during the patent prosecution process, including signing the formal USPTO declaration

**Inventors named in §2 sign here:**

| Inventor | Signature | Date |
|---|---|---|
| Anas Abu Ehmade (Founder) | _pending_ | _pending_ |
| Head of AI Development | _pending_ | _pending_ |

**Witness (Copyright & IP Officer):**

| Officer | Signature | Date |
|---|---|---|
| Senior Copyright & IP Officer | _pending_ | _pending_ |

---

## 12 · Outside counsel handoff

When this disclosure is finalised, the Copyright Officer engages outside patent counsel from the firm's pre-qualified shortlist (maintained as a private file). The handoff package contains:

- This disclosure document (the present file)
- Source-code excerpts demonstrating the resolver, selection logic, JSON salvager, and compact-prompt fallback (with proprietary-coefficient redactions where appropriate)
- Architecture diagrams produced by the Head of AI Development
- Sales artefacts 40 (Hybrid AI Sheet) for marketing-narrative consistency
- The trademark-filing brief for parallel Osolix™ registration

Outside counsel produces the formal USPTO-format application within 30 days of receiving the package.

---

## 13 · Trade-secret disclosure boundary

The following elements of the implementation are deliberately **omitted** from the public-facing patent claims to preserve them as trade secrets:

- The exact small-model name pattern matched by the compact-prompt fallback
- The specific text of the compact prompt
- The proprietary coefficients in the offline-model confidence-score calibration
- The audit-log row format details beyond the schema headers

These remain Osolix LLC's confidential trade secrets and are protected separately under the employee + contractor IP-assignment agreements + the master subscription agreement IP indemnity clauses.

---

## 14 · Sign-off

By approving this provisional patent disclosure, the Founder + Head of AI Development + Copyright Officer commit to:

- Funding the §10.1 budget for Year 1
- Inventor cooperation with outside counsel during prosecution
- Maintaining the trade-secret boundary in §13
- Annual review of the patent portfolio at the Q3 IP committee meeting

**Latest sign-off:** _pending — 2026-05-02_ · Targeted US provisional filing: 2026-Q3 mid-quarter
