# fileAI Docs --- # Welcome to fileAI Source: https://docs.file.ai/index ## On this page * [About us](#about-us) * [Quick links](#quick-links) Get Started # Welcome to fileAI Copy pageCopy page The foundation for reliable, scalable AI workflows. Copy pageCopy page ## [​ ](#about-us) About us fileAI is an AI-native data preparation platform that helps you turn unstructured and fragmented files into clean, structured data you can trust. Our platform is built to handle the heavy lifting of document processing — from reading and understanding different file types, to enriching content with context, to validating accuracy with citations. With fileAI, you can: * Upload and process almost any file format, including spreadsheets, PDFs, text documents, and images * Automatically structure and align your data to schemas with zero-shot, automated generation — no manual formatting required * Enrich data across files to surface insights and reduce duplication * Ensure data accuracy and compliance through built-in validation fileAI processes hundreds of millions of files each year for global enterprises across industries, helping teams reduce manual work, improve decision-making, and scale automation with confidence. ## [​ ](#quick-links) Quick links ## Quickstart Guide Get your workspace set up in minutes with our quickstart guide ## AI Schemas Zero-shot data schemas for structured output ## Folders Intelligent and intuitive file management ## Account Set up and manage your fileAI account ## MCP Give your AI Agents access to fileAI ## APIs Simplify data preparation and workflows with APIs ## Support Contact support for account & billing questions [ Quickstart ](/docs-user-guide/get-started/getting-started) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Quickstart Guide Source: https://docs.file.ai/docs-user-guide/get-started/getting-started ## On this page * [Step 1: Create an account, workspace, and project](#step-1-create-an-account-workspace-and-project) * [Step 2: Set up your workspace](#step-2-set-up-your-workspace) * [Step 3: Create tags and upload files](#step-3-create-tags-and-upload-files) Get Started # Quickstart Guide Copy pageCopy page Get started on fileAI Copy pageCopy page ## [​ ](#step-1-create-an-account-workspace-and-project) Step 1: Create an account, workspace, and project ## [​ ](#step-2-set-up-your-workspace) Step 2: Set up your workspace ## [​ ](#step-3-create-tags-and-upload-files) Step 3: Create tags and upload files [ Introduction ](/)[ Best Practices ](/docs-user-guide/get-started/best-practice) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # fileAI best practices Source: https://docs.file.ai/docs-user-guide/get-started/best-practice ## On this page * [File types](#file-types) * [How to upload a file](#how-to-upload-a-file) * [How to export a file](#how-to-export-a-file) Get Started # fileAI best practices Copy pageCopy page Guidelines to get the best output from fileAI Copy pageCopy page ## [​ ](#file-types) File types We support the following file types: Excel & CSV (XLS, XLSX, CSV) Ideal for structured tabular data. For best results, use **single-sheet files** with a **header row** and consistent columns. PDF **PDF (PDF):** Works best with **digitally generated PDFs** (not scanned) and clear layout. Tables and text should be selectable. Word & Text (DOCX, TXT) Structured, paragraph-based content is supported. Headings and sections help improve understanding. Images (JPG, JPEG, PNG, TIF, TIFF, HEIC, HEIF) Supported for scanned documents or forms. Ensure **clear resolution and readable text** for optimal processing. ## [​ ](#how-to-upload-a-file) **How to upload a file** 1. Go to the Drive. 2. Drag and drop your file or click the black button at the right panel to Upload. Wait for the processing bar to complete. File status will update (e.g. _Transform_, _Processed, Review_). ## [​ ](#how-to-export-a-file) How to export a file 1. Select the file you’d like to export. 2. Go to the **AI Schema** tab to review the data extracted from your file. 3. Click the **Download** button dropdown at the upper right. 4. Choose from the available export options: XLSX, XLS, CSV, JSON, PDF, XERO\_CSV * **Download as XERO\_CSV** is only for bank statements and generates a Xero-compatible CSV file. If you try to use **XERO\_CSV** for a non–bank statement file, an error message will appear. * All other export options work with every file type, so you can select whichever suits your needs. 5. After you choose which format to download your extracted data in, a progress bar will appear. 6. When the download finishes, you’ll see a notification, and the file will be saved in your **Downloads** folder. [ Quickstart ](/docs-user-guide/get-started/getting-started)[ FAQs ](/docs-user-guide/get-started/faqs) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Frequently Asked Questions Source: https://docs.file.ai/docs-user-guide/get-started/faqs Get Started # Frequently Asked Questions Copy pageCopy page Your questions, answered Copy pageCopy page How does fileAI handle my data security and privacy? At fileAI, protecting your data is our highest priority. Every file is encrypted in transit (TLS) and at rest (AES-256), with strict role-based access controls and audit logs in place. We operate on ISO 27001-certified infrastructure, and our controls are independently audited for SOC 2 Type 2. fileAI is also fully GDPR and HIPAA compliant.For organizations with advanced security needs, we offer private cloud and on-premise deployment options to ensure full data isolation.You can explore our certifications and controls in detail at our [trust portal](https://trust.file.ai), powered by Vanta. What types of files and formats can I use with fileAI? We support most file formats, including: JPEG, PNG, TIF, TIFF, HEIC, HEIF, PDF, CSV, XLSX, XLS, DOC, DOCX, ZIP, WEBP. For any file formats not listed, please reach out to our team. Can I integrate fileAI with my existing tools? Yes. fileAI connects with the tools you already use, so you can bring all your business knowledge into one place. Today, you can integrate fileAI with: * Google Drive * Google Sheets * Dropbox * OneDrive * Email * Xero * [API](https://docs.file.ai/docs-api/api-intro) & [**MCP**](https://docs.file.ai/docs-mcp/get-started-mcp) (Model Context Protocol) With **MCP**, fileAI becomes instantly usable inside environments like **Cursor and Claude**, and can work alongside other MCP servers such as **Notion**. This means your agents can use OCR, classification, and field extraction without custom integrations, and chain fileAI into multi-step workflows across your stack.We are also building more integrations based on customer feedback, with NetSuite and SharePoint coming soon. How accurate and reliable is fileAI's outputs? fileAI is designed to provide accurate answers and summaries from your files, but as with any AI system, there may be occasional gaps or errors. To give you confidence in the results, we’ve built **Citation & Validation** directly into the platform. * **Citations:** fileAI shows exactly where extracted values or answers come from in the original document, so you can instantly verify the source without manual cross-checking. * **Validation:** In addition to citations, our system applies reasoning and quality checks to flag potential issues. For example, if values conflict, contain OCR errors, or need human review, fileAI highlights them with a clear status (e.g., _verified_ or _needs review_). * **Transparency:** This combination of AI-driven extraction plus human-verifiable citations ensures traceability and auditability, making fileAI particularly reliable for compliance-sensitive use cases. Accuracy improves further when files are structured and up to date, and we provide feedback tools so you can refine results over time. This means you are always in control — able to trust answers while verifying them when it matters most. What does the pricing model look like, and when will I hit limits? fileAI uses a straightforward pay-as-you-go model. You only pay for what you process, with no limits on users, workspaces, or storage. One billable page equals either one document page or one spreadsheet sheet. When you sign up, you’ll also receive a $5 credit to try the system. Billing can be monthly or annual.**Plan differences** * **Self-serve**: Best for individuals and small teams. You get everything needed to start quickly — all file formats, core AI schemas, and unlimited users — with simple pay-per-page pricing. * **Enterprise**: Designed for complex, high-volume automation. This plan includes advanced options such as custom-trained models, unlimited schemas, private cloud or on-premise deployment, workflow orchestration, and a guaranteed uptime SLA. For the latest details, visit our [pricing page](https://file.ai/pricing). For any unanswered billing or account issues, please reach out to [support@file.ai](mailto:support@file.ai) [ Best Practices ](/docs-user-guide/get-started/best-practice)[ AI Schemas ](/docs-user-guide/core/ai-schemas) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # AI Schemas Source: https://docs.file.ai/docs-user-guide/core/ai-schemas ## On this page * [AI Schemas](#ai-schemas) * [How to edit a field](#how-to-edit-a-field) * [How to edit an existing AI Schema](#how-to-edit-an-existing-ai-schema) * [Schema Locking](#schema-locking) * [What is schema locking?](#what-is-schema-locking) * [How can my project benefit from schema locking?](#how-can-my-project-benefit-from-schema-locking) * [How to enable schema locking](#how-to-enable-schema-locking) Core # AI Schemas Copy pageCopy page AI Schemas let you produce a deterministic data schema, from a single file, cross-file or with online data fetch capabilities Copy pageCopy page ## [​ ](#ai-schemas) AI Schemas Power your workflows and AI agents with our robust data schema collection. Our schemas will: * Automatically classify files accurately, without renaming or hallucinations * Enforce a consistent, deterministic data schema for extracting metadata from each file * Enrich file metadata with online data - _cross-file enrichment coming soon_ Saving of your schema(s) can be done conveniently within our UI. You can let our system auto-classify and suggest a schema, or create your own schemas at your convenience. ### [​ ](#how-to-edit-a-field) How to edit a field 1. Open a processed file. 2. Click on the **AI Schema** button in the upper right panel and edit each tab according to your preferences.  3. You may also click the Sparkle icon next to the field to enter a prompt describing how you’d like the field to be captured then click Run and Save to apply the changes.  4. You can also mark a field as thumbs up or thumbs down if they are correct or incorrect.  ✏️These changes will eventually become an AI Schema and apply to future files uploaded.  ### [​ ](#how-to-edit-an-existing-ai-schema) How to edit an existing AI Schema Tweak your existing AI Schema by following these steps:  1. Click **File Intelligence** on the left panel below Drive. 2. Go to the **Extraction Schema** tab. 3. Select and click the schema you want to edit. 4. From there, you can: 1. Add a new table or field. 2. Remove existing ones by clicking the trash icon.  3. Hide the data by clicking the eye icon.  4. Use the pencil icon to make changes to existing fields.  5. Once you’re done, click the Save AI schema on the upper right side.  6. Your changes will apply to newly processed files. ⚠️ Schema edits are version-controlled. Older files use previous versions unless re-processed. ## [​ ](#schema-locking) Schema Locking ### [​ ](#what-is-schema-locking) **What is schema locking?** Schema locking is a feature that prevents the system from creating new file types. When it’s enabled, your project is limited to the schemas you’ve already saved, so new documents can only be classified into those existing schemas. It ensures files are processed according to your locked schemas, instead of creating new ones. ### [​ ](#how-can-my-project-benefit-from-schema-locking) **How can my project benefit from schema locking?** Having a fixed schema applied to all documents in your Project can help:  * **Consistency** – Documents are only classified into your schemas, reducing unexpected or incorrect new file types. * **Quality Control** – Prevents the system to create too many unnecessary or duplicate schemas. * **Simplified Review** – Easier to manage and validate results since all documents fit into a fixed set of classes. * **Stronger Workflows** – Useful in compliance-heavy or structured environments where only certain document types are allowed.  ### [​ ](#how-to-enable-schema-locking) **How to enable schema locking** 1. **Edit your AI Schema** Open the file and update the AI Schema if needed (view the [help center guide](https://docs.file.ai/docs-user-guide/core/ai-schemas#how-to-edit-an-existing-ai-schema) above for detailed instructions). 2. **Verify the AI Schema** Go to the upper right of the page, open the _Verify AI Schema_ dropdown, and confirm. _Tip: You can also use the AI Schema Editor in the upper right corner to rename your AI Schema if needed._ 3. **View confirmed schemas** In the left panel, click _**File Intelligence**_ then go to the _**Extraction Schema**_ tab to see the list of all schemas you’ve verified. 4. **Enable Schema Locking** In the _**Extraction Schema**_ tab, go to the upper right and turn on the _S**chema Locked**_ toggle. From now on, uploaded files will only be classified according to the schemas you’ve locked. [ FAQs ](/docs-user-guide/get-started/faqs)[ Folders ](/docs-user-guide/core/folders) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Folders Source: https://docs.file.ai/docs-user-guide/core/folders ## On this page * [Capabilities](#capabilities) * [Folder Categories](#folder-categories)  * [How to Use Folders](#how-to-use-folders) * [1\. Creating a New Folder](#1-creating-a-new-folder) * [2\. Uploading Files into Folders](#2-uploading-files-into-folders) * [3\. Renaming a Folder](#3-renaming-a-folder) Core # Folders Copy pageCopy page Folders are digital containers used to group related files together in a meaningful way. They support both manual organization and future logic-based automation. Copy pageCopy page This guide provides a clear overview of how to effectively use **Folders** within the platform. Folders are a core organizing feature that help users manage, group, and navigate large volumes of files with ease. * * * ## [​ ](#capabilities) **Capabilities** ## Reference-based file structure * Files live in a **central AI Drive** * Folders reference files, they do **not duplicate** * Files can be referenced in **multiple folders** at once ## Folder levels supported * Parent folder * Subfolder * * * ## [​ ](#folder-categories) Folder Categories  Manual Folders & Subfolders * These folders are created **manually** by users, and are fully user-controlled with standard file management options * **You can:** * ✅ Upload files * ✅ Remove files * ✅ Delete the folder or subfolder Auto-Created Subfolders These folders are\*\*automatically generated \*\* based on rules set in the parent folder (e.g., via route + match + compare logic). and are used to organize files that meet certain criteria.**You cannot:** * ❌ Manually upload or add files * ❌ Remove files * ❌ Delete the subfolder directly 🔄 **Note**: These subfolders are dynamic — if deleted (via parent), they will regenerate automatically as long as the parent folder’s rules remain. Parent Folders with Rules **These folders are manually created** by users, and include automation rules that organize files into subfolders.**You can:** * ✅ Delete the parent folder * When deleted, all of its auto-created subfolders are also removed. * The files inside remain available in their original location (e.g., Drive / All Files). Auto-Created Parent Folders (Export Folders) These folders are**automatically created** when connecting to integrations (e.g., Xero, Odoo), and are system-generated based on data sources.**You can:** * ✅ Delete the parent folder ⚠️ Behavior may differ slightly from standard folders — updates coming to align functionality. ## [​ ](#how-to-use-folders) **How to Use Folders** ### [​ ](#1-creating-a-new-folder) **1\. Creating a New Folder** 1. Go to **Drive** 2. Click **“+”** next to Drive tab 3. Enter a name in the dialog box 4. Click **Create** * * * ### [​ ](#2-uploading-files-into-folders) **2\. Uploading Files into Folders** * **Drag & drop** into folder * Use **Upload Modal** * * * ### [​ ](#3-renaming-a-folder) **3\. Renaming a Folder** 1. Locate the folder 2. Enter the folder view 3. Click the **pencil icon** 4. Enter the new name 5. Press **Confirm** * * * [ AI Schemas ](/docs-user-guide/core/ai-schemas)[ Citations & Validations ](/docs-user-guide/core/citations-and-validations) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Citations & Validations Source: https://docs.file.ai/docs-user-guide/core/citations-and-validations ## On this page * [About](#about) * [How do I enable Citations and Validations in my Project?](#how-do-i-enable-citations-and-validations-in-my-project)  * [Utilizing Citations and Validations](#utilizing-citations-and-validations) * [How do I view Citations and Validations in my file?](#how-do-i-view-citations-and-validations-in-my-file) * [How do I interpret validation icons (✅ vs ⚠️)?](#how-do-i-interpret-validation-icons--vs--) * [How do I manually validate extracted values if a field needs review (⚠️)?](#how-do-i-manually-validate-extracted-values-if-a-field-needs-review--)  Core # Citations & Validations Copy pageCopy page Citations & Validations address one of the biggest concerns in AI-driven automation: trust. Copy pageCopy page ## [​ ](#about) About **Citations and validations** make it easier to trust AI results. fileAI shows you the original source of information in a file and flags whether the extracted data has been validated. This gives you confidence in the output, improves accuracy, and helps you meet compliance requirements. Citations and Validations are designed specifically for PDFs, JPEGs, tabular data within PDFs and JPEGs, and files under 100 page. That said, there are a few things Citations & Validations can’t handle yet: * They don’t support Excel or CSV files. * They don’t work on PDFs longer than 100 pages. Reach out to your account manager or [support@file.ai](mailto:support@file.ai) to surface this feature on your workspace. ### [​ ](#how-do-i-enable-citations-and-validations-in-my-project) How do I enable Citations and Validations in my Project?  If you’ve requested to surface the feature, you’ll see the option to enable them in your Project settings. 1. Go to your Project Settings. 2. Look for the Citations & Validations toggle. 3. Enable the feature. 4. Save your settings. 5. Once enabled, all extracted values will include citations and, when applicable, validation results. If you don’t see this option, and have requested to surface it, contact your admin or support to upgrade. ## [​ ](#utilizing-citations-and-validations) Utilizing Citations and Validations ### [​ ](#how-do-i-view-citations-and-validations-in-my-file) **How do I view Citations and Validations in my file?** 1. Open the file once it’s done processing. 2. Switch to the AI Schema tab to view structured data fields automatically extracted from the file. 3. Click the validation status icon (✅ vs ⚠️) to view the validation status. ### [​ ](#how-do-i-interpret-validation-icons--vs--) **How do I interpret validation icons (✅ vs ⚠️)?** Validation results are shown with simple icons: * ✅ Green Check: The extracted value has been validated successfully against the original document. * ⚠️ Warning Sign: The value could not be validated automatically and needs manual review. These icons help you quickly spot fields that need attention. ### [​ ](#how-do-i-manually-validate-extracted-values-if-a-field-needs-review--) How do I manually validate extracted values if a field needs review (⚠️)?  Sometimes, the system requires human review to confirm accuracy. In these cases, do the following:  1. Click the ⚠️validation icon 2. Hover to the right panel to see the Citation and Validation results.  3. Review the comments to check why the field needs review. 4. You can manually type the correct data based on your records. 5. Finalize the review by marking as validated, confirming it is correct.  6. The icon will update to a ✅green checkmark once done.  [ Folders ](/docs-user-guide/core/folders)[ File Intelligence ](/docs-user-guide/core/file-intelligence) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # File Intelligence Source: https://docs.file.ai/docs-user-guide/core/file-intelligence ## On this page * [Where to Find File Intelligence](#where-to-find-file-intelligence) * [General Settings](#general-settings) * [Extraction Schema](#extraction-schema) * [Notification](#notification) Core # File Intelligence Copy pageCopy page Copy pageCopy page The **File Intelligence** tab allows you to control how documents are processed by the AI. From this tab, you can configure how files are split, what language the system prioritizes, which parsing model is used, and how documents are automatically classified. All settings configured in this section act as your **project’s default settings** for document processing. At this time, **File Intelligence settings apply across the entire project** and cannot be customized at the folder or file level. When a project is created, all settings in File Intelligence are preconfigured with recommended defaults. You only need to adjust them if you have specific processing requirements. ## [​ ](#where-to-find-file-intelligence) **Where to Find File Intelligence** 1. Go to **AI Drive** 2. Select **File Intelligence** from the left navigation menu.  The File Intelligence tab contains three sections: * **General Settings** – Configure how documents are processed * **Extraction Schema** – Define the structure used for AI data extraction * **Notification** – Manage notification preferences # [​ ](#general-settings) **General Settings** General Settings control the **behavior for how documents are processed by the AI system** based on your project requirements. If no changes are made, the system uses the following default configuration.  **File Splitting** Determines how multi-page documents are processed. * **Do not split (Default)** The entire document is processed as a single file. * **Smart splitting** Each page is automatically separated and processed as an individual document. **Excel / CSV File Splitting** Controls how spreadsheet files are handled during processing. * **Do not split (Default)** The spreadsheet is processed as one file. * **Smart splitting** Large spreadsheets are automatically divided into smaller sections to improve processing performance. **File Target Language** Specifies the language the AI should prioritize when interpreting document content. * **Auto-detect (Default)**  The system automatically detects the language in the document. * **Supported languages**  You can manually select a specific language if needed. Tip: In most cases, **Auto-detect provides the best results**. **AI Parsing Model** Determines which AI model reads and interprets document content. **Available options include** * \*\*Beethoven\_ENG\_GP25 (Default) \*\* Optimized for general document processing and works well for most use cases. * \*\*Beethoven\_Direct\_Form\_Filling GP2.5 (DTFF) \*\* Designed for specialized document workflows and may not be required for typical processing. * Other available AI OCR models.  **ML/AI-based Classification** Automatically categorizes documents using a pre-trained machine learning model during processing. This helps organize documents and improve downstream automation. * \*\*Active — System default  \*\* This setting enables automatic document classification across all processed files. * * * # [​ ](#extraction-schema) **Extraction Schema** The Extraction Schema tab provides access to all **AI Schemas** configured in your project. To learn more about creating and managing AI Schemas, refer to the [**AI Schema documentation**](https://docs.file.ai/docs-user-guide/core/ai-schemas). * * * # [​ ](#notification) **Notification** The **Notification** tab allows you to manage alerts related to document processing. To learn more about creating and managing Notifications, refer to the [Notification documentation](https://dashboard.mintlify.com/fileai/fileai/editor/main/~/4fc628ca-5a40-489f-b0b6-f135b1daaa4f).  [ Citations & Validations ](/docs-user-guide/core/citations-and-validations)[ Notifications ](/docs-user-guide/core/notifications) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Notifications Source: https://docs.file.ai/docs-user-guide/core/notifications Core # Notifications Copy pageCopy page Copy pageCopy page The **Notifications** feature allows you to receive automated email alerts based on file activity and workflow conditions. Instead of manually monitoring files, you can create rules that trigger notifications when specific criteria are met. **This ensures the right people are informed at the right time.** When a file is processed, the system checks it against your notification rules. If the conditions are met, an email is sent to the selected recipients. Each rule is also linked to a verified AI Schema, which defines the conditions, data fields, and recipients used for the notification.  The key benefits here is this helps streamline your workflows by reducing manual monitoring and automating notifications based on your document data. * **Automated alerts** – Set rules once and receive notifications automatically * **Accurate triggers** – Notifications are based on actual document data * **Flexible delivery** – Send alerts in real-time or on a schedule * **Dynamic recipients** – Notify the right people using data from the document * **Audit visibility** – Track when and why notifications were sent [ File Intelligence ](/docs-user-guide/core/file-intelligence)[ Account Set Up ](/docs-user-guide/account/account1) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Account Set Up Source: https://docs.file.ai/docs-user-guide/account/account1 ## On this page * [How to Sign Up](#how-to-sign-up) * [How to Invite and Manage Members](#how-to-invite-and-manage-members) Account # Account Set Up Copy pageCopy page How to set up your fileAI account Copy pageCopy page ### [​ ](#how-to-sign-up) **How to Sign Up** 1. Go to our [homepage](https://orion.file.ai/) and click **Sign Up**. 2. Enter your work email and set a password. 3. Check your email for the verification link. 4. Once verified, you’ll be redirected to your dashboard. 🔐 **Tip:** Use a company email to ensure proper workspace setup. * * * ### [​ ](#how-to-invite-and-manage-members) **How to Invite and Manage Members** 1. Go to your **Workspace name** dropdown at the top left corner > **Settings** > **Members**. 2. Click **Invite Members** and enter their email. 3. Select the project then assign a role within the project and workspace: Admin or Participant. 4. To manage permissions of existing members, click the 3-dot menu next to each user. 👥 User roles provide different access to schema editing, API usage, and billing. [ Notifications ](/docs-user-guide/core/notifications)[ Billing ](/docs-user-guide/account/billing) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Billing & Support Source: https://docs.file.ai/docs-user-guide/account/billing ## On this page * [How to Navigate to the Billing Page](#how-to-navigate-to-the-billing-page) * [What happens when I run out of credits?](#what-happens-when-i-run-out-of-credits) * [How can I upgrade to an Enterprise Plan?](#how-can-i-upgrade-to-an-enterprise-plan) * [How to Properly Raise a Bug](#how-to-properly-raise-a-bug) Account # Billing & Support Copy pageCopy page Copy pageCopy page fileAI V2 uses a **pay-as-you-go model** to give users flexibility and control. When you sign up, you receive a **$5 credit** to help you get started. This allows you to explore the platform without committing to a full subscription right away. **Note:** The $5 credit expires **30 days** after you sign up. Be sure to use it before then! * * * ### [​ ](#how-to-navigate-to-the-billing-page) **How to Navigate to the Billing Page** 1. Click your workspace icon at the top left of the page. 2. Go to **Settings > Billing**. 3. You’ll see: * Current plan/credits * Current storage * Current file type within the workspace * Current import/export number of the file * Usage breakdown (Upload date/time, Processing time/Cost/Etc.) 💬 For billing questions, click **Contact Support** on this page. * * * ### [​ ](#what-happens-when-i-run-out-of-credits) **What happens when I run out of credits?** If you use up your credits, you can continue using V2 by: * Making a **one-time top-up**, or * Setting up **automatic top-ups** to keep your balance active. You can manage your credit settings anytime from your **Billing** section in your account dashboard. * * * ### [​ ](#how-can-i-upgrade-to-an-enterprise-plan) **How can I upgrade to an Enterprise Plan?** Need more features or a higher credit limit? You can upgrade to an **Enterprise Plan**, which offers: * Custom credit options * Dedicated support * Enterprise-level features To upgrade, simply contact our sales team via this[form](https://share.hsforms.com/1zX6UE9ewRd6bYONZkENH2g4m622) and someone will reach out to discuss your needs. * * * ### [​ ](#how-to-properly-raise-a-bug) **How to Properly Raise a Bug** 1. Click conversation bubbles icon to contact Support team 2. Fill out: * Description of the issue * Your project/workspace name * Related document links/IDs (if available) * Screenshots or screen recording 3. Submit and wait for the support agent to response 🐞 Use specific filenames or URLs when possible for faster resolution. [ Account Set Up ](/docs-user-guide/account/account1) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # About the fileAI API Source: https://docs.file.ai/docs-api/api-intro ## On this page * [Key Features](#key-features) * [Base URL and instances](#base-url-and-instances) * [Switching between instances](#switching-between-instances) * [Prerequisites for using the fileAI API](#prerequisites-for-using-the-fileai-api) * [How to get an API token](#how-to-get-an-api-token) * [How to authorize an API token](#how-to-authorize-an-api-token) fileAI API # About the fileAI API Copy pageCopy page Explore the core capabilities of the fileAI API Copy pageCopy page Our fileAI API is organized around REST and uses standard HTTP methods (POST, GET, PATCH) to interact with resource objects. It uses JSON as the primary data format, and all requests should have a Content-Type of application/json. ## [​ ](#key-features) Key Features These API endpoints are the building blocks for all your file ingestion needs. They can be used individually or in conjunction to create powerful end-to-end workflows. ## File Upload & Management Upload, retrieve and manage files in various formats in fileAI (includes sources, citations, location data and more) ## Smart File Processing Proprietary AI OCR with cross-file extraction and custom AI schemas ## Dynamic AI Schemas Auto-generated data structures that adapt to your files, fetch cross-file data (coming soon) and online data all in one workflow ## Enterprise Ready Secure, scalable, and built for production workloads (SOCII, ISO27001) ## Easy Integration RESTful API with callback URLs for real-time updates Our fileAI API is organized around REST and uses standard HTTP methods (POST, GET, PATCH) to interact with resource objects. It uses JSON as the primary data format, and all requests should have a Content-Type of application/json. ## [​ ](#base-url-and-instances) Base URL and instances The default base URL is: ``` https://api.orion.file.ai/prod/v1 ``` Unless you have been told otherwise, this is the base URL to use, and it is the one every example in these docs is written against. fileAI also runs instance-specific hosts. An instance serves exactly the same API — identical paths, methods, parameters, request bodies and response schemas — and differs **only in the hostname**: Instance Region Base URL Default — `https://api.orion.file.ai/prod/v1` `au` Australia `https://api.orion.au.file.ai/prod/v1` `sg` Singapore `https://api.orion.sg.file.ai/prod/v1` `jp` Japan `https://api.orion.jp.file.ai/prod/v1` So every endpoint in this reference is reached at: ``` https://api.orion.file.ai/prod/v1/ ← default https://api.orion..file.ai/prod/v1/ ``` Not sure whether your workspace is on an instance, or which one? Ask your account manager or contact [support@file.ai](mailto:support@file.ai). ## [​ ](#switching-between-instances) Switching between instances Because only the hostname changes, switching instances never means rewriting a request. Swap the host and everything else stays byte-for-byte identical. * In your code * In a curl command * In the API playground Keep the base URL in one constant or environment variable and build every request path from it. Switching instances is then a one-line config change, not a search-and-replace. ``` # Default export FILEAI_BASE_URL="https://api.orion.file.ai/prod/v1" # Singapore instance export FILEAI_BASE_URL="https://api.orion.sg.file.ai/prod/v1" ``` ``` curl -X GET "$FILEAI_BASE_URL/files" \ -H "x-api-key: YOUR_API_KEY" ``` Replace the host, leave the path, query string, headers and body untouched. ``` # Default curl -X GET "https://api.orion.file.ai/prod/v1/files" \ -H "x-api-key: YOUR_API_KEY" # Same call, Japan instance curl -X GET "https://api.orion.jp.file.ai/prod/v1/files" \ -H "x-api-key: YOUR_API_KEY" ``` Every endpoint page in this reference has an interactive playground. Use the **server** dropdown above the request to switch from the default host to `https://api.orion.{instance}.file.ai`, then pick `au`, `sg` or `jp` from the **instance** field. The sample request updates to match, so you can copy a snippet that already points at the right host. A workspace belongs to exactly one instance, and an API key only authenticates against the host that issued it. Sending a request to a different instance fails authentication even though the endpoint exists there — so switch the base URL and the API key together. ## [​ ](#prerequisites-for-using-the-fileai-api) Prerequisites for using the fileAI API Before using our API, please ensure you complete the following prerequisites: ## You must have a fileAI account To use our API endpoints, you need to sign up or login [here](https://orion.file.ai/en/sign-up). Both self-serve and enterprise accounts are supported. ## You must have an API Key After creating your fileAI account, you can generate your API Key. ## You must verify an AI Schema fileAI suggests extraction and data fetch schemas. Confirm or edit these in the UI to call them directly via MCP. ### [​ ](#how-to-get-an-api-token) **How to get an API token** 1. Go to **API Keys** 2. Click **Create API Key** 3. Name the token, set expiration and permission 4. Click **Generate API Key** 5. Token will be shown, click **Copy & Close** Note: Only admins can create API keys. * * * ### [​ ](#how-to-authorize-an-api-token) **How to authorize an API token** 1. Once you have your API key, go to the API documentation page: `https://api.orion.file.ai/prod/v1/docs/swagger-ui`On an instance-specific host, swap the hostname — for example `https://api.orion.sg.file.ai/prod/v1/docs/swagger-ui`. See [Base URL and instances](#base-url-and-instances). 2. Click the Authorize button on the right of this page (the green Authorize button with lock icon) 3. Enter your API Key under Value 4. Click Authorize to start making authenticated requests directly from the documentation If you need any further assistance, please reach out to our team at [support@file.ai](mailto:support@file.ai) [ Quick Start ](/docs-api/quick-start) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Quick Start Source: https://docs.file.ai/docs-api/quick-start fileAI API # Quick Start Copy pageCopy page Get up and running quickly with simple steps to authenticate, upload files, and extract data using the fileAI API. Copy pageCopy page 1 Test Your Connection ## API Reference ![](https://files.readme.io/cd64c5f60bdf6faf1b1cd803398edaede6dbf27dd720d0055638d6a55acad43f-carbon_2.png) **Success Response:** `{"message": "Public API is up and running!"}` 2 Upload Your First File **Request an upload URL:** ## API Reference ![](https://files.readme.io/75242f660a44712bb88563ebd407f317c394cd4f3467d275bff5815dcedaf9c3-carbon_6.png) \*_Upload your file to the returned presigned URL, then check your results:_ ## API Reference ![](https://files.readme.io/b70d981b76c37b1c04e44026bd32ad89c076b19041a36e50539910fb12869990-carbon_7.png) That’s it! Your file is now processed and structured data is extracted automatically. [ Introduction ](/docs-api/api-intro)[ Upload ](/docs-api/upload) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Upload Source: https://docs.file.ai/docs-api/upload ## On this page * [Upload Overview](#upload-overview) * [When to use the Upload Endpoint](#when-to-use-the-upload-endpoint) * [Key Features](#key-features) Key Functions # Upload Copy pageCopy page Understanding fileAI’s Upload endpoint Copy pageCopy page ## [​ ](#upload-overview) Upload Overview Our upload endpoint allows you to send files directly to fileAI. ## [​ ](#when-to-use-the-upload-endpoint) When to use the Upload Endpoint ## Submitting New Files for Processing Use the [upload endpoint](https://developers.file.ai/reference/publicapicontroller_uploadfilerequest)t to send new files (PDFs, images, or text-based files) to fileAI for data extraction and schema application ## Automating File Intake Integrate the upload endpoint into your system to automate the ingestion of files from users, third-party tools, or internal workflows ## Triggering AI Schema Extraction Uploading a file initiates the AI-driven extraction process, which applies the appropriate schema and prepares the data for validation or export ## [​ ](#key-features) Key Features 1. Currently Supported File Types include: Image Formats * PNG * JPEG/JPG * GIF * TIFF PDF * PDF (Portable Document Format) Spreadsheets * CSV (Comma Separated Values) * XLSX (Excel Open XML) * XLSM (Excel with Macros) * XLS (Excel Binary) 2. Direct uploads have a **50MB** file size limit 3. Our platform can automatically split bulk files and identify individual files as part of the upload process. [ Quick Start ](/docs-api/quick-start)[ AI OCR Models ](/docs-api/api-ai-ocr-models) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # AI OCR models Source: https://docs.file.ai/docs-api/api-ai-ocr-models ## On this page * [How to Use](#how-to-use) * [List of AI OCR models available](#list-of-ai-ocr-models-available) Key Functions # AI OCR models Copy pageCopy page We’re excited to bring advanced AI Optical Character Recognition (OCR) capabilities to your applications through our latest AI models to support a variety of use cases. Copy pageCopy page ### [​ ](#how-to-use) **How to Use** Our AI OCR capabilities are available via the [/POST endpoint.](https://developers.file.ai/reference/publicapicontroller_uploadfilerequest#/) * Simply send a file to receive structured text output — _with exact bounding box coordinates for text position coming soon_ * Fetch additional data from online sources via your custom AI schema — _with cross-file retrieval coming soon_ ### [​ ](#list-of-ai-ocr-models-available) **List of AI OCR models available** English Models: * Beethoven\_ENG\_O5.6 - OpenAI v6 * Beethoven\_ENG\_G5.5 - Gemini v5 * Beethoven\_ENG\_GP25 - Gemini Pro 2.5 * Beethoven\_ENG\_GP25.1 - Gemini Pro 2.5 v1 * Beethoven\_ENG\_GP25.2 - Gemini Pro 2.5 PDF * Beethoven\_ENG\_GP3 - Gemini Pro 3 * Beethoven\_CUS\_O5.1 - Custom OpenAI v8 * Beethoven\_CUS\_O5.2 - Custom Gemini v13 * Unified (google-document-ai-ocr-gemini-v10) - Unified model Chinese Models: * Beethoven\_ZH\_O5.9 - Chinese OpenAI v9 Japanese Models: * Beethoven\_JP\_O5.3 - Japanese OpenAI v3 * Beethoven\_JP\_G5.4 - Japanese Gemini fine-tuned Thai Models: * Beethoven\_TH\_O5.1 - Thai OpenAI v1 [ Upload ](/docs-api/upload)[ AI Schemas ](/docs-api/api-ai-schema) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # AI Schemas Source: https://docs.file.ai/docs-api/api-ai-schema ## On this page * [AI Data Schemas - Your Powerful Workflow Enabler](#ai-data-schemas-your-powerful-workflow-enabler) Key Functions # AI Schemas Copy pageCopy page AI Schemas let you produce a deterministic data schema, from a single file, cross-file or with online data fetch capabilities Copy pageCopy page ## [​ ](#ai-data-schemas-your-powerful-workflow-enabler) **AI Data Schemas - Your Powerful Workflow Enabler** Power your workflows and AI agents with our robust data schema collection. Our schemas will: * Automatically classify files accurately, without renaming or hallucinations * Enforce a consistent, deterministic data schema for extracting metadata from each file * Enrich file metadata with online data - _cross-file enrichment coming soon_ [ AI OCR Models ](/docs-api/api-ai-ocr-models)[ Supported Data Types ](/docs-api/api-supported-data-types) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Supported Data Types Source: https://docs.file.ai/docs-api/api-supported-data-types ## On this page * [File Naming Convention](#file-naming-convention) * [Supported Data Types](#supported-data-types) * [1\. text](#1-text) * [2\. yes/no](#2-yes%2Fno) * [3\. number](#3-number) * [4\. date-time](#4-date-time) * [5\. enum](#5-enum) Key Functions # Supported Data Types Copy pageCopy page Our API accepts the following data types for inputs and outputs. Please ensure that your data conforms to one of these types for compatibility. Copy pageCopy page ## [​ ](#file-naming-convention) **File Naming Convention** All field names follow the **snake\_case** format: `_` * Use lowercase letters only. * Words are separated by underscores (`_`). * Examples: * `received_date` * `user_name` * `is_active` * `status_enum` This consistent naming ensures readability and clarity across all API fields. ## [​ ](#supported-data-types) **Supported Data Types** ### [​ ](#1-text) 1\. `text` * **Description**: A string of characters. * **Example**: `"Hello, world!"` ### [​ ](#2-yes/no) 2\. `yes/no` * **Description**: A boolean value representing a binary choice. * **Accepted Values**: `true` / `false` (or `yes` / `no`, depending on context) * **Example**: `true` ### [​ ](#3-number) 3\. `number` * **Description**: Any numeric value, including integers and floats. * **Example**: `42`, `3.14` ### [​ ](#4-date-time) 4\. `date-time` * **Description**: A date and time string in [ISO 8601](https://en.wikipedia.org/wiki/ISO_8601) format. * **Format**: `YYYY-MM-DDTHH:MM:SSZ` * **Example**: `"2025-06-12T14:30:00Z"` ### [​ ](#5-enum) 5\. `enum` * **Description**: A predefined set of string values. Use this type when the value must be one of a known set, where status can be one of: “pending”, “approved”, “rejected” * **Example**: JSON ``` { "status": "pending" } ``` [ AI Schemas ](/docs-api/api-ai-schema)[ Extraction Prompts ](/docs-api/api-extraction-prompts) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Extraction prompts Source: https://docs.file.ai/docs-api/api-extraction-prompts ## On this page * [1\. Be Explicit and Specific](#1-be-explicit-and-specific) * [2\. Match Field Names Exactly](#2-match-field-names-exactly) * [3\. Use Clear Instructions](#3-use-clear-instructions) * [4\. Set Expectations for Format](#4-set-expectations-for-format) * [5\. Provide Context When Necessary](#5-provide-context-when-necessary) * [6\. Use Examples (if supported)](#6-use-examples-if-supported) Key Functions # Extraction prompts Copy pageCopy page Best practices for writing effective prompts to extract structured data in fileAI Copy pageCopy page ### [​ ](#1-be-explicit-and-specific) **1\. Be Explicit and Specific** * Instead of “Extract important details”, use “Extract the invoice\_number, total\_amount, and due\_date from the text.” ### [​ ](#2-match-field-names-exactly) **2\. Match Field Names Exactly** * Use field names as defined in your schema, e.g., customer\_name, not name or client. ### [​ ](#3-use-clear-instructions) **3\. Use Clear Instructions** * Example: “Extract the delivery\_status\_enum as one of: shipped, pending, delayed.” ### [​ ](#4-set-expectations-for-format) **4\. Set Expectations for Format** * For dates: “Extract received\_date in the format YYYY-MM-DDTHH:MM:SSZ.” * For yes/no: “Is the payment confirmed? Respond with true or false.” ### [​ ](#5-provide-context-when-necessary) **5\. Provide Context When Necessary** * If the field is ambiguous, clarify in the prompt * Example:“Extract the contract\_date\_time (the date the contract was signed, not created).“ ### [​ ](#6-use-examples-if-supported) **6\. Use Examples (if supported)** * If your system supports few-shot prompting, include examples of expected input/output. [ Supported Data Types ](/docs-api/api-supported-data-types)[ Schema Locking ](/docs-api/api-schema-locking) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Schema Locking Source: https://docs.file.ai/docs-api/api-schema-locking ## On this page * [Overview](#overview) * [How It Works](#how-it-works) * [Visual Comparison](#visual-comparison) * [Scenario: Uploading a Parking Ticket](#scenario-uploading-a-parking-ticket) * [Schema Locking OFF](#schema-locking-off) * [Schema Locking ON](#schema-locking-on) * [When to Use Schema Locking](#when-to-use-schema-locking) * [Enable Schema Locking When:](#enable-schema-locking-when) * [Keep Schema Locking Disabled When:](#keep-schema-locking-disabled-when) * [Best Practices](#best-practices) * [Getting Started](#getting-started) * [Maintenance](#maintenance) * [Team Workflows](#team-workflows) Key Functions # Schema Locking Copy pageCopy page Description of your new file. Copy pageCopy page ## [​ ](#overview) **Overview** Schema Locking is a configuration setting that controls how your document processing system handles new or unfamiliar document types. This feature gives you precise control over file classification behavior, ensuring consistency and preventing unwanted schema proliferation. ## [​ ](#how-it-works) **How It Works** Your system uses intelligent document classification that learns from the documents you upload and confirm. As you process more documents, the system becomes better at recognizing patterns and automatically categorizing similar files. ## [​ ](#visual-comparison) **Visual Comparison** ![Screenshot 2025-08-18 at 4.43.35 PM.png](/images/Screenshot2025-08-18at4.43.35PM.png "Screenshot 2025-08-18 at 4.43.35 PM.png") ### [​ ](#scenario-uploading-a-parking-ticket) **Scenario: Uploading a Parking Ticket** **Your Current Schemas:** * Invoice * Bank Statement * Job Contract #### [​ ](#schema-locking-off) **Schema Locking OFF** ``` Parking Ticket Upload ↓ 🤖 "This doesn't match any existing schemas" ↓ 📝 Creates: "Unknown" classification ↓ 💡 Suggests: "Parking Ticket" as new schema ↓ ✅ You confirm → Now you have 4 schemas ``` #### [​ ](#schema-locking-on) **Schema Locking ON** ``` Parking Ticket Upload ↓ 🤖 "Must choose from existing schemas only" ↓ 🎯 Forced classification → "Invoice" (best match) ↓ 🔒 No new schema created ``` ## [​ ](#when-to-use-schema-locking) **When to Use Schema Locking** ### [​ ](#enable-schema-locking-when) **Enable Schema Locking When:** **You have a fixed document workflow** * Your business processes a specific set of document types * You want to prevent schema drift over time * Multiple team members upload documents and you need consistency **You want to avoid misclassifications** * Preventing edge cases where similar-looking documents create unnecessary schemas * Maintaining a clean, organized schema structure * Ensuring all documents fit into predefined categories **Enterprise environments** * Compliance requirements mandate specific document categories * You need predictable classification behavior * Integration with downstream systems expects consistent schema names ### [​ ](#keep-schema-locking-disabled-when) **Keep Schema Locking Disabled When:** **You’re in discovery mode** * Still learning what document types your organization processes * Want the system to identify new document patterns automatically * Regularly encounter new document formats **Flexible document processing** * Your document types vary significantly over time * You want maximum automation in schema detection * Prefer to review and approve new schemas as they’re discovered ## [​ ](#best-practices) **Best Practices** ### [​ ](#getting-started) **Getting Started** 1. **Begin with Schema Locking OFF** to discover your document patterns 2. **Monitor new schema suggestions** 3. **Enable Schema Locking** once you’ve identified your core document types ### [​ ](#maintenance) **Maintenance** * **Review classification accuracy** periodically when Schema Locking is enabled * **Temporarily disable** Schema Locking when you know new document types are coming * **Clean up schemas** before enabling locking to avoid forcing documents into inappropriate categories ### [​ ](#team-workflows) **Team Workflows** * **Document your confirmed schemas** for team members * **Set clear guidelines** on when to temporarily disable locking * **Regular audits** ensure locked schemas still meet business needs * * * [ Extraction Prompts ](/docs-api/api-extraction-prompts)[ Processing Callback ](/docs-api/api-processing-callback) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Processing callback Source: https://docs.file.ai/docs-api/api-processing-callback ## On this page * [Overview](#overview) * [🚀 Integration Guide](#-integration-guide) * [1\. Set up your webhook endpoint](#1-set-up-your-webhook-endpoint) * [2\. Configure webhook URL](#2-configure-webhook-url) * [3\. Handle events in your application](#3-handle-events-in-your-application) * [Core Fields](#core-fields) * [🔄 Event Types](#-event-types) * [1\. File Uploaded](#1-file-uploaded) * [2\. Processing Started](#2-processing-started) * [3\. Fraud Detection](#3-fraud-detection) * [4\. First Page Analysis](#4-first-page-analysis) * [5\. OCR Processing](#5-ocr-processing) * [6\. OCR Evaluation](#6-ocr-evaluation) * [7\. Split & Classify](#7-split-%26-classify) * [8\. Processing Finished](#8-processing-finished) Key Functions # Processing callback Copy pageCopy page Real-time updates for your file processing pipeline through secure webhook callbacks. Copy pageCopy page ## [​ ](#overview) **Overview** Our file processing system sends webhook events to notify your application about the status of uploaded files as they progress through our processing pipeline. Each event contains detailed information about the current processing step and file status. **📋 Quick Facts** * **Method**: `POST` * **Content-Type**: `application/json` * **Retries**: Up to 5 attempts with exponential backoff * **Timeout**: 30 seconds * * * ## [​ ](#-integration-guide) **🚀 Integration Guide** ### [​ ](#1-set-up-your-webhook-endpoint) **1\. Set up your webhook endpoint** Create an HTTP endpoint that accepts POST requests and returns a `200` status code to acknowledge receipt of the webhook. ### [​ ](#2-configure-webhook-url) **2\. Configure webhook URL** Set your webhook URL in the dashboard or contact support to configure your callback endpoint. ### [​ ](#3-handle-events-in-your-application) **3\. Handle events in your application** Your endpoint will receive JSON payloads for each processing step. Parse the `status`, `step`, `uploadId`, and `fileIds` fields to track file processing progress. ``` --- ## 📊 Processing Workflow Your files progress through our automated processing pipeline: ``` 📁 File Upload ↓ ⚡ Processing Started ↓ 🛡️ Fraud Detection ↓ 📄 First Page Analysis ↓ 🔍 OCR Processing ↓ ✅ OCR Evaluation ↓ 📑 Split & Classify ↓ 🎉 Processing Finished ```` --- ## 📋 Event Reference ### Base Event Structure All webhook events share this common structure: ```json { "status": "completed", "step": "processing_started", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:30:00Z", "fileIds": "b765bd00-181e-4737-b4d9-004dadd6bd45" } ```` #### [​ ](#core-fields) **Core Fields** **Field** **Type** **Description** `status` `string` Current status: `waiting`, `completed`, `failed` `step` `string` Processing step identifier `uploadId` `string` Unique identifier for the upload session `timestamp` `string` ISO 8601 timestamp when event occurred `fileIds` `string` Available in final steps - processed file IDs * * * ## [​ ](#-event-types) **🔄 Event Types** ### [​ ](#1-file-uploaded) **1\. File Uploaded** **Step**: `file_uploaded` **When**: File successfully uploaded to our system JSON ``` { "status": "completed", "step": "file_uploaded", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:30:00Z" } ``` **Next Step**: `processing_started` * * * ### [​ ](#2-processing-started) **2\. Processing Started** **Step**: `processing_started` **When**: File enters processing queue JSON ``` { "status": "waiting", "step": "processing_started", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:30:15Z" } ``` JSON ``` { "status": "completed", "step": "processing_started", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:30:30Z" } ``` **Status Flow**: `waiting` → `completed` **Next Step**: `fraud_detection` * * * ### [​ ](#3-fraud-detection) **3\. Fraud Detection** **Step**: `fraud_detection` **When**: Security and fraud checks completed JSON ``` { "status": "completed", "step": "fraud_detection", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:31:00Z" } ``` **Next Step**: `first_page_analysis` * * * ### [​ ](#4-first-page-analysis) **4\. First Page Analysis** **Step**: `first_page_analysis` **When**: Initial document analysis completed JSON ``` { "status": "completed", "step": "first_page_analysis", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:31:30Z" } ``` **Next Step**: `beethoven_ocr` * * * ### [​ ](#5-ocr-processing) **5\. OCR Processing** **Step**: `beethoven_ocr` **When**: Optical Character Recognition completed JSON ``` { "status": "completed", "step": "beethoven_ocr", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:32:15Z" } ``` **Next Step**: `ocr_evaluation` * * * ### [​ ](#6-ocr-evaluation) **6\. OCR Evaluation** **Step**: `ocr_evaluation` **When**: OCR quality assessment completed JSON ``` { "status": "completed", "step": "ocr_evaluation", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "timestamp": "2025-06-11T10:32:45Z" } ``` **Next Step**: `split_classify` * * * ### [​ ](#7-split-&-classify) **7\. Split & Classify** **Step**: `split_classify` **When**: Document splitting and classification completed JSON ``` { "status": "completed", "step": "split_classify", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "fileIds": ["b765bd00-181e-4737-b4d9-004dadd6bd45"], "timestamp": "2025-06-11T10:33:15Z" } ``` **📌 Important**: First event where `fileIds` becomes available **Next Step**: `processing_finished` * * * ### [​ ](#8-processing-finished) **8\. Processing Finished** **Step**: `processing_finished` **When**: All processing completed successfully JSON ``` { "status": "completed", "step": "processing_finished", "uploadId": "e9818d17-e5be-4d57-bdc8-5a40b4f6f4e1", "fileIds": ["b765bd00-181e-4737-b4d9-004dadd6bd45"], "timestamp": "2025-06-11T10:33:30Z" } ``` **🎉 Final Step**: Your files are ready for download/use [ Schema Locking ](/docs-api/api-schema-locking)[ Incremental Loading (CDC) ](/docs-api/api-incremental-data-loading) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Incremental Data Loading Source: https://docs.file.ai/docs-api/api-incremental-data-loading ## On this page * [Overview](#overview) * [Timestamp Filters](#timestamp-filters) * [Timestamp Contract](#timestamp-contract) * [Watermark Semantics](#watermark-semantics) * [Inclusive Comparison (>=)](#inclusive-comparison-%3E%3D) * [Deduplication Requirement](#deduplication-requirement) * [Ordering & Pagination](#ordering-%26-pagination) * [Default Sort Order](#default-sort-order) * [Recommended Sort for CDC](#recommended-sort-for-cdc) * [Pagination Parameters](#pagination-parameters) * [Pagination Stability Warning](#pagination-stability-warning) * [What Triggers updatedAt?](#what-triggers-updatedat) * [Delete Handling](#delete-handling) * [Tracking Deletions](#tracking-deletions) * [Best Practices](#best-practices) * [Snowflake Integration Example](#snowflake-integration-example) * [Response Schema](#response-schema) * [Quick Reference](#quick-reference) * [Need Help?](#need-help) Key Functions # Incremental Data Loading Copy pageCopy page Sync files incrementally to data warehouses using Change Data Capture patterns with the GET /files endpoint. Copy pageCopy page ## [​ ](#overview) Overview The `GET /files` endpoint supports **incremental data loading** through timestamp-based filters. This enables efficient Change Data Capture (CDC) workflows for syncing to Snowflake, BigQuery, or other data warehouses. Instead of fetching all files on every sync, use `updatedAfter` or `createdAfter` to retrieve only what’s changed since your last sync. * * * ## [​ ](#timestamp-filters) Timestamp Filters ## updatedAfter Filters files updated **on or after** the specified timestamp. `GET /files?updatedAfter=2025-01-01T00:00:00.000Z` ## createdAfter Filters files created **on or after** the specified timestamp. `GET /files?createdAfter=2025-01-01T00:00:00.000Z` ### [​ ](#timestamp-contract) Timestamp Contract Property Value **Format** ISO 8601 **Timezone** UTC **Precision** Milliseconds **Example** `2025-01-01T00:00:00.000Z` All timestamps in the response (`createdAt`, `updatedAt`) are returned in UTC with millisecond precision. * * * ## [​ ](#watermark-semantics) Watermark Semantics ### [​ ](#inclusive-comparison-\>=) Inclusive Comparison (`>=`) The `updatedAfter` and `createdAfter` filters use **inclusive** comparison. This means: * A file with `updatedAt = 2025-01-01T12:00:00.000Z` **will be returned** when querying with `updatedAfter=2025-01-01T12:00:00.000Z` ### [​ ](#deduplication-requirement) Deduplication Requirement Because of inclusive comparison, **clients may receive duplicate records** across paginated requests or subsequent syncs. You must deduplicate using `fileId + updatedAt`. **Example deduplication in Snowflake:** ``` MERGE INTO target_table t USING staging_table s ON t.file_id = s.file_id WHEN MATCHED AND s.updated_at > t.updated_at THEN UPDATE SET ... WHEN NOT MATCHED THEN INSERT ... ``` * * * ## [​ ](#ordering-&-pagination) Ordering & Pagination ### [​ ](#default-sort-order) Default Sort Order Parameter Default Value `sortBy` `createdAt` `sortOrder` `ASC` ### [​ ](#recommended-sort-for-cdc) Recommended Sort for CDC When using `updatedAfter` for incremental loading, always sort by `updatedAt` in ascending order for reliable watermark tracking. ``` GET /files?updatedAfter=2025-01-01T00:00:00.000Z&sortBy=updatedAt&sortOrder=ASC ``` ### [​ ](#pagination-parameters) Pagination Parameters The API uses **offset-based pagination** with `page` and `limit` parameters. Parameter Default Maximum `page` 1 — `limit` 100 100 **Example request:** ``` GET /files?updatedAfter=2025-01-01T00:00:00.000Z&sortBy=updatedAt&sortOrder=ASC&page=1&limit=100 ``` ### [​ ](#pagination-stability-warning) Pagination Stability Warning **Offset-based pagination may produce inconsistent results** if records are updated during pagination. You may miss records or see duplicates. **Mitigation strategies:** 1. Use smaller time windows for `updatedAfter` 2. Always deduplicate by `fileId + updatedAt` 3. Re-sync periodically with a larger time window to catch missed records * * * ## [​ ](#what-triggers-updatedat) What Triggers `updatedAt`? The `updatedAt` timestamp changes when any of these events occur: Event Updates `updatedAt` File metadata changes ✅ Yes Schema/form field updates ✅ Yes Document status changes ✅ Yes Form filling/reprocessing ✅ Yes Tag modifications ✅ Yes Contact/vendor updates ✅ Yes Workflow step changes ✅ Yes Classification changes ✅ Yes OCR reprocessing ✅ Yes * * * ## [​ ](#delete-handling) Delete Handling **Important**: Deleted files are **excluded** from the `GET /files` response and do **not** appear in `updatedAfter` queries. This endpoint provides **insert/update only** — not full CDC. ### [​ ](#tracking-deletions) Tracking Deletions If you need to detect deleted files: 1 Option A: Periodic Full Sync Do a full sync periodically and compare with your existing data to detect missing records. 2 Option B: Deletion Events Endpoint Contact the fileAI team about a dedicated deletion events endpoint (roadmap item). * * * ## [​ ](#best-practices) Best Practices 1\. Store Your Watermark After each successful sync, store the **maximum `updatedAt`** value from the batch: ``` # Pseudocode last_sync_watermark = max(record['updatedAt'] for record in batch) save_watermark(last_sync_watermark) ``` 2\. Use Recommended Query Parameters Always include sorting parameters for reliable CDC: ``` GET /files?updatedAfter={watermark}&sortBy=updatedAt&sortOrder=ASC&limit=100 ``` 3\. Handle Pagination Completely Paginate through **all pages** before updating your watermark: ``` # Pseudocode page = 1 all_records = [] while True: response = get_files(updatedAfter=watermark, page=page, limit=100) all_records.extend(response['files']) if len(response['files']) < 100: break page += 1 # Only update watermark after all pages are processed if all_records: new_watermark = max(r['updatedAt'] for r in all_records) save_watermark(new_watermark) ``` 4\. Deduplicate Before Loading Always deduplicate records by `fileId + updatedAt` before inserting into your data warehouse. 5\. Schedule Periodic Full Syncs Run full syncs (e.g., weekly) to: * Catch any records missed due to pagination issues * Detect deleted records by comparing with existing data * * * ## [​ ](#snowflake-integration-example) Snowflake Integration Example * Initial Load * Incremental Load ``` -- First sync: load all files COPY INTO raw_files FROM ( SELECT $1:fileId, $1:fileName, $1:updatedAt, ... FROM @fileai_stage/files.json ); ``` ``` -- Create staging table for incremental data CREATE TEMPORARY TABLE staging_files AS SELECT * FROM raw_files WHERE 1=0; -- Load incremental data COPY INTO staging_files FROM ( SELECT $1:fileId, $1:fileName, $1:updatedAt, ... FROM @fileai_stage/incremental_files.json ); -- Merge with deduplication MERGE INTO raw_files t USING ( SELECT * FROM staging_files QUALIFY ROW_NUMBER() OVER ( PARTITION BY file_id ORDER BY updated_at DESC ) = 1 ) s ON t.file_id = s.file_id WHEN MATCHED AND s.updated_at > t.updated_at THEN UPDATE SET file_name = s.file_name, updated_at = s.updated_at, ... WHEN NOT MATCHED THEN INSERT (file_id, file_name, updated_at, ...) VALUES (s.file_id, s.file_name, s.updated_at, ...); ``` * * * ## [​ ](#response-schema) Response Schema ``` { "files": [ { "fileId": "53d6a0b1-2a8d-4ed9-9e6a-ceaef7ca3908", "fileName": "invoice.pdf", "fileType": "application/pdf", "fileSize": 261928, "fileStoragePath": "path/to/file.pdf", "fileHash": "be3ef5fbb21e31c1281300b23b1c918a8ba54427c799aea21865f68d5efd01b7", "uploadId": "f2538513-f0b9-4aa8-9c57-bc0a85c77de6", "status": "processed", "currency": "USD", "summary": "Invoice from Example Corp", "duplicateToFileId": null, "referenceId": "1e70a5e860", "isDuplicate": false, "isEmbedded": false, "schemaId": "6835aca6030a79ffaabca742", "fileClass": "Invoice", "fileContactId": "6835aca2281d9ed1bab90b11", "fileContactName": "Example Company", "createdAt": "2025-05-27T12:14:24.258Z", "updatedAt": "2025-05-27T14:30:00.123Z" } ], "count": 1, "currentPage": 1 } ``` * * * ## [​ ](#quick-reference) Quick Reference Feature Behavior Filter semantics `>=` (inclusive) Timestamp format ISO 8601, UTC, milliseconds Pagination Offset-based (`page`, `limit`) Default sort `createdAt ASC` Recommended sort for CDC `updatedAt ASC` Delete visibility ❌ Not visible (insert/update only) Deduplication required ✅ Yes, by `fileId + updatedAt` * * * ## [​ ](#need-help) Need Help? ## Contact Support Reach out to the fileAI engineering team for questions about CDC implementation or to request new features. [ Processing Callback ](/docs-api/api-processing-callback)[ Pagination ](/docs-api/api-pagination) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Pagination Source: https://docs.file.ai/docs-api/api-pagination ## On this page * [Overview](#overview) * [Endpoints that support cursor pagination](#endpoints-that-support-cursor-pagination) * [Cursor pagination](#cursor-pagination) * [Parameters](#parameters) * [Pagination object](#pagination-object) * [Legacy page/limit pagination](#legacy-page%2Flimit-pagination) * [Choosing a mode](#choosing-a-mode) Key Functions # Pagination Copy pageCopy page Page through large result sets with cursor pagination, or keep using the legacy page/limit parameters. Copy pageCopy page ## [​ ](#overview) Overview List endpoints support two pagination modes. Which one you get is decided by a single parameter: **send `cursor` and you are in cursor mode, omit it and you stay in legacy `page`/`limit` mode**. Cursor pagination is the recommended mode for anything that walks a full result set. It is stable under concurrent writes, and its cost does not grow as you move deeper into the results. Legacy `page`/`limit` pagination is still fully supported and unchanged. No existing integration needs to be updated. ## [​ ](#endpoints-that-support-cursor-pagination) Endpoints that support cursor pagination Endpoint Legacy response key Cursor mode notes [`GET /files`](/api-reference/endpoint/get-all-files) `files` — [`GET /schemas`](/api-reference/endpoint/get-all-schemas) `schemas` — [`GET /directories`](/api-reference/endpoint/get-all-directories) `directories` `data` is a **flat** list of directories — no `subfolders` nesting. `cursor`, `page`, `limit`, `sortBy`, and `sortOrder` come from one shared pagination parameter set, so `cursor` is also _listed_ on [`GET /file-type`](/api-reference/endpoint/get-file-types), [`GET /schemas/file-type`](/api-reference/endpoint/get-file-type-schemas), and [`GET /files/{fileId}/values`](/api-reference/endpoint/get-file-schema-values-by-fileids). Those endpoints do not implement cursor mode yet: sending `cursor` there is accepted but ignored, and you still get the legacy `page`/`limit` response. Use `page`/`limit` on them. ## [​ ](#cursor-pagination) Cursor pagination Examples use the default base URL, `https://api.orion.file.ai/prod/v1`. If your workspace is on an instance-specific host, swap the hostname and change nothing else — see [Switching between instances](/docs-api/api-intro#switching-between-instances). 1 Request the first page Send `cursor` with an **empty value** — the parameter’s presence is what selects cursor mode. ``` curl -X GET "https://api.orion.file.ai/prod/v1/files?cursor=&limit=50" \ -H "x-api-key: YOUR_API_KEY" ``` 2 Read the pagination block The response is wrapped in `data` plus a `pagination` object. ``` { "data": [{ "id": "649e2d2d2d2d2d2d2d2d2d2d" }], "pagination": { "limit": 50, "nextCursor": "eyJjcmVhdGVkQXQiOiIyMDI1LTA3LTMxVDA5OjE0OjIyLjEwMFoifQ==", "hasMore": true } } ``` 3 Follow nextCursor until hasMore is false Pass the previous `nextCursor` value back as `cursor`. When `hasMore` is `false`, `nextCursor` is `null` and you have reached the end. ``` curl -X GET "https://api.orion.file.ai/prod/v1/files?cursor=eyJjcmVhdGVkQXQiOiIyMDI1LTA3LTMxVDA5OjE0OjIyLjEwMFoifQ%3D%3D&limit=50" \ -H "x-api-key: YOUR_API_KEY" ``` Cursor values are opaque and URL-encode them before use — they are base64 and can contain `=` characters. ### [​ ](#parameters) Parameters Parameter Type Required Default Description cursor string No — Opaque cursor. Present (even empty) selects cursor mode; omitted keeps legacy. limit number No 50 Page size in cursor mode. Maximum 100. sortBy string No `createdAt` In cursor mode, one of `createdAt`, `updatedAt`, `_id`. sortOrder string No `ASC` `ASC` or `DESC`. ### [​ ](#pagination-object) Pagination object Field Type Description limit number Page size actually applied (default 50, max 100). nextCursor string | null Cursor for the next page; `null` when `hasMore` is `false`. hasMore boolean Whether more results exist after this page. A cursor is bound to the `sortBy`, `sortOrder`, and filter values it was minted with. Changing any of them mid-walk invalidates the cursor and returns a `422`. Start a new walk instead. ## [​ ](#legacy-page/limit-pagination) Legacy page/limit pagination Omit `cursor` entirely and the response keeps its original shape: a named array (see the table above), plus `count` and `currentPage`. ``` curl -X GET "https://api.orion.file.ai/prod/v1/files?page=1&limit=100" \ -H "x-api-key: YOUR_API_KEY" ``` ``` { "files": [{ "id": "649e2d2d2d2d2d2d2d2d2d2d" }], "count": 1350, "currentPage": 1 } ``` Parameter Type Required Default Description page number No 1 Page number. limit number No 100 Items per page. sortBy string No `createdAt` Field to sort by. sortOrder string No `ASC` Sort direction (`ASC` or `DESC`) ## [​ ](#choosing-a-mode) Choosing a mode ## Use cursor pagination Walking a full result set, syncing to a warehouse, or paging through data that is being written to concurrently. ## Use page/limit You need a total `count`, jump-to-page behaviour, or you have an existing integration that already works. Legacy mode gives you a total `count`; cursor mode does not — computing an exact total is what makes deep offset paging expensive. Use `hasMore` to drive your loop instead of a total. [ Incremental Loading (CDC) ](/docs-api/api-incremental-data-loading)[ Idempotency ](/docs-api/api-idempotency) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Idempotent requests Source: https://docs.file.ai/docs-api/api-idempotency ## On this page * [Overview](#overview) * [Using it](#using-it) * [Outcomes](#outcomes) * [Cached failures](#cached-failures) * [Endpoints](#endpoints) * [Accepted but ignored](#accepted-but-ignored) Key Functions # Idempotent requests Copy pageCopy page Send an Idempotency-Key header so a retried write cannot be applied twice. Copy pageCopy page ## [​ ](#overview) Overview Network timeouts leave you guessing: did the upload land, or not? Retrying is safe only if the API can recognise the retry. Every write endpoint accepts an optional `Idempotency-Key` request header for exactly that. The first request with a given key does the work and its response is stored for **24 hours**. A retry with the same key **and the same body** replays that stored response instead of doing the work again, and carries the header `Idempotent-Replayed: true`. The header is optional everywhere. Existing integrations keep working unchanged — omit it and requests behave exactly as before. ## [​ ](#using-it) Using it Examples use the default base URL, `https://api.orion.file.ai/prod/v1`. If your workspace is on an instance-specific host, swap the hostname and change nothing else — see [Switching between instances](/docs-api/api-intro#switching-between-instances). 1 Generate a key per logical operation Any opaque string up to 255 characters. A UUIDv4 is recommended. Reuse the _same_ key for every retry of the same operation, and a fresh key for a new operation. 2 Send it as a request header ``` curl -X POST "https://api.orion.file.ai/prod/v1/directories" \ -H "x-api-key: YOUR_API_KEY" \ -H "Idempotency-Key: 0f1c8e03-978e-40d5-bc93-6894a57f9324" \ -H "Content-Type: application/json" \ -d '{"name": "Q3 Invoices"}' ``` 3 Retry with the same key on failure On a timeout, `429`, or `5xx`, retry with the identical key and body. If the original request had already succeeded, you get its response back with `Idempotent-Replayed: true` and nothing is created twice. ## [​ ](#outcomes) Outcomes Situation Result First request with this key Executes normally. Response stored for 24h. Same key, same body, original finished `200`\-series replay of the stored response, `Idempotent-Replayed: true`. Same key, same body, original **still running** `409` `IDEMPOTENT_REQUEST_IN_PROGRESS`. Same key, **different** body `422` `IDEMPOTENCY_KEY_REUSED`. Key older than 24h Treated as a new key — the request executes again. No `Idempotency-Key` header No idempotency handling at all. “Same body” is determined by a fingerprint over the method, path, workspace, and a canonical form of the request body — key ordering and whitespace do not matter, so you can resend a re-serialized body safely. ### [​ ](#cached-failures) Cached failures A terminal failure is cached and replayed too — a `403`, `404`, or business `409`/`422` will not be re-executed under the same key. Two classes are deliberately **not** cached, so a retry gets a fresh attempt: * Transient statuses (`408`, `425`, `429`) — a later retry may legitimately succeed. * `VALIDATION_FAILED` and `BAD_REQUEST` — a corrected body has a different fingerprint and must not be locked out for 24 hours. ## [​ ](#endpoints) Endpoints `Idempotency-Key` is honoured on every write endpoint: ## Files `POST /files/upload` · `POST /files/upload/multipart` · `POST /files/upload/multipart/complete` · `DELETE /files` · `PATCH /files/{fileId}` ## Schemas & values `PATCH /files/schema` · `PATCH /files/{fileId}/values` · `PATCH /files/schema/rerun` · `PATCH /schemas` · `POST /files/schema/export` · `POST /files/schema/import` ## File types `PATCH /file-type/approve` · `PATCH /file-type/rename` ## Workspace `POST /directories` · `PATCH /project/setting` ### [​ ](#accepted-but-ignored) Accepted but ignored Three `POST` endpoints are semantically reads. They accept the header for client uniformity but never replay: every call executes fresh, with no 24h window, no `Idempotent-Replayed` header, and no key-reuse `422`. Endpoint Why `POST /query/execute` No side effects, and a later run can legitimately return different rows. `POST /files/preview` Presigned URLs expire in an hour — a 24h replay would hand back dead URLs. `POST /files/preview/batch` Same as above. Idempotency fails **open**. If the backing store is briefly unreachable the request still proceeds, but non-idempotently — it is a resilience aid, not a transactional guarantee. Keep your own retry bookkeeping for operations where a duplicate would be costly. [ Pagination ](/docs-api/api-pagination)[ Error Responses ](/docs-api/api-errors) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Error responses Source: https://docs.file.ai/docs-api/api-errors ## On this page * [Overview](#overview) * [Fields](#fields) * [Field-level violations](#field-level-violations) * [Error codes](#error-codes) * [Legacy fields](#legacy-fields) * [Handling errors](#handling-errors) Key Functions # Error responses Copy pageCopy page Every fileAI API error is an RFC 9457 problem document with a stable machine-readable code and a request correlation id. Copy pageCopy page ## [​ ](#overview) Overview Every error response from the API is an [RFC 9457](https://www.rfc-editor.org/rfc/rfc9457) problem document, served with `Content-Type: application/problem+json`. ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/prod/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName" } ] } ``` Branch your error handling on **`code`**, not on `detail` or `title`. `code` is a stable public contract — new codes may be added, but existing ones are never renamed or repurposed. `detail` is human-facing prose and can change at any time. ## [​ ](#fields) Fields Field Type Description type string URI identifying the problem type, e.g. `https://errors.file.ai/validation-failed`. title string Stable, human-readable summary of the problem type. status number HTTP status code, repeated in the body. detail string Human-readable explanation specific to this occurrence. instance string URI reference for this occurrence — the full request path, including the `/prod` prefix. code string Stable machine-readable error code — branch on this. requestId string Correlation id for the request. Quote it in support requests. errors array Field-level violations. Populated on validation errors, `[]` otherwise. retryAfter number Seconds until you may retry. Present on `429` only. ### [​ ](#field-level-violations) Field-level violations Each entry in `errors` describes one offending field: Field Type Description code string Violation code, e.g. `REQUIRED`. detail string What was wrong, e.g. `must not be empty`. pointer string JSON Pointer to the offending body field, e.g. `#/fileName`. parameter string Name of the offending query or header parameter, e.g. `limit`. `pointer` is set for body fields and `parameter` for query/header parameters — an entry carries whichever one applies. ## [​ ](#error-codes) Error codes Code Status Title When it happens `BAD_REQUEST` 400 Bad request The request could not be processed as sent. `UNAUTHENTICATED` 401 Unauthenticated The `x-api-key` header is missing or the key is invalid. `FORBIDDEN` 403 Forbidden The key is inactive, or lacks access to the resource. `READONLY_KEY` 403 Read-only API key A read-only key was used on a write endpoint. `NOT_FOUND` 404 Not found The referenced resource does not exist in this workspace. `IDEMPOTENT_REQUEST_IN_PROGRESS` 409 Idempotent request in progress An earlier request with the same `Idempotency-Key` is still running. `VALIDATION_FAILED` 422 Validation failed One or more parameters or body fields are invalid. See `errors[]`. `IDEMPOTENCY_KEY_REUSED` 422 Idempotency key reused The same `Idempotency-Key` was reused with a different request body. `RATE_LIMITED` 429 Too many requests The rate limit was exceeded. Wait `retryAfter` seconds. `INTERNAL_ERROR` 500 Internal server error Unexpected server-side failure. Safe to retry. ## [​ ](#legacy-fields) Legacy fields Problem documents also carry three deprecated keys for backwards compatibility with integrations written before RFC 9457 responses were introduced: Legacy field Use instead `message` `detail` / `title` `error` `title` `statusCode` `status` The legacy keys are deprecated and will be removed in a future version. Migrate to `detail`, `title`, and `status`. ## [​ ](#handling-errors) Handling errors Retrying safely `429` and `500` are safe to retry. On `429`, wait `retryAfter` seconds; for `500`, back off exponentially. Send an [`Idempotency-Key`](/docs-api/api-idempotency) on write endpoints so a retry cannot apply the same change twice. Do not retry `400`, `401`, `403`, `404`, and `422` describe a problem with the request itself. Retrying without changing anything returns the same error. Reporting a problem Include the `requestId` and `code` when contacting [support@file.ai](mailto:support@file.ai) — `requestId` correlates your call to our server-side logs. [ Idempotency ](/docs-api/api-idempotency)[ Get all files ](/api-reference/endpoint/get-all-files) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Get all files Source: https://docs.file.ai/api-reference/endpoint/get-all-files get /prod/v1/files Get all files of the current organization and workspace. This endpoint is paginated. Use the `page` and `limit` query parameters to control pagination. Alternatively, pass the `cursor` query parameter (empty for the first page) to use cursor pagination: the response envelope becomes `{ data, pagination: { limit, nextCursor, hasMore } }` and pages are stable under concurrent writes. Follow `nextCursor` until `hasMore` is false; keep the same sort and filter parameters for every page of a walk. Use the optional `updatedAfter` query parameter to filter files that have been updated on or after the specified timestamp (ISO 8601 format). Use the optional `createdAfter` query parameter to filter files that have been created on or after the specified timestamp (ISO 8601 format). These filters are useful for batch processes that need to fetch only files modified or created since the last sync. Use the optional `directoryId` query parameter to return only files that belong to the specified directory (or sub-directory). Files in nested descendants are not included; pass the sub-directory id directly to fetch its files. --- # Update file export status Source: https://docs.file.ai/api-reference/endpoint/update-file-export-status patch /prod/v1/files/{fileId} This endpoint allows integrated systems to update the export status of a file. Use this endpoint for pull-based integrations where integrated systems need to report back: - Whether a file was successfully exported to the integrated system - Any error messages or status information from the integrated system The export status and message will be visible in the file list UI. --- # Delete multiple files Source: https://docs.file.ai/api-reference/endpoint/delete-multiple-files Delete multiple files cURL ``` curl --request DELETE \ --url https://api.orion.file.ai/prod/v1/files \ --header 'Content-Type: application/json' \ --header 'x-api-key: ' \ --data ' { "fileIds": [ "8123612d-187c-44eb-8ac6-214e68506b7e", "8123612d-187c-44eb-8ac6-214e68506b7f" ] } ' ``` ``` import requests url = "https://api.orion.file.ai/prod/v1/files" payload = { "fileIds": ["8123612d-187c-44eb-8ac6-214e68506b7e", "8123612d-187c-44eb-8ac6-214e68506b7f"] } headers = { "x-api-key": "", "Content-Type": "application/json" } response = requests.delete(url, json=payload, headers=headers) print(response.text) ``` ``` const options = { method: 'DELETE', headers: {'x-api-key': '', 'Content-Type': 'application/json'}, body: JSON.stringify({ fileIds: ['8123612d-187c-44eb-8ac6-214e68506b7e', '8123612d-187c-44eb-8ac6-214e68506b7f'] }) }; fetch('https://api.orion.file.ai/prod/v1/files', options) .then(res => res.json()) .then(res => console.log(res)) .catch(err => console.error(err)); ``` ``` "https://api.orion.file.ai/prod/v1/files", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 30, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "DELETE", CURLOPT_POSTFIELDS => json_encode([ 'fileIds' => [ '8123612d-187c-44eb-8ac6-214e68506b7e', '8123612d-187c-44eb-8ac6-214e68506b7f' ] ]), CURLOPT_HTTPHEADER => [ "Content-Type: application/json", "x-api-key: " ], ]); $response = curl_exec($curl); $err = curl_error($curl); curl_close($curl); if ($err) { echo "cURL Error #:" . $err; } else { echo $response; } ``` ``` package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.orion.file.ai/prod/v1/files" payload := strings.NewReader("{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}") req, _ := http.NewRequest("DELETE", url, payload) req.Header.Add("x-api-key", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(string(body)) } ``` ``` HttpResponse response = Unirest.delete("https://api.orion.file.ai/prod/v1/files") .header("x-api-key", "") .header("Content-Type", "application/json") .body("{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}") .asString(); ``` ``` require 'uri' require 'net/http' url = URI("https://api.orion.file.ai/prod/v1/files") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Delete.new(url) request["x-api-key"] = '' request["Content-Type"] = 'application/json' request.body = "{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}" response = http.request(request) puts response.read_body ``` 200 401 403 404 422 429 500 ``` { "success": true } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "Invalid fileIds.", "instance": "/v1/{route}", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [], "message": "Invalid fileIds.", "error": "Unprocessable Entity", "statusCode": 422 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` Endpoints # Delete multiple files Copy pageCopy page This endpoint is used to delete multiple files. The `fileIds` field is used to specify the files that should be deleted. The files that are specified in the `fileIds` field will be deleted. If a files is not specified in the `fileIds` field, the files will not be deleted. Copy pageCopy page DELETE https://api.orion.file.aihttps://api.orion.{instance}.file.ai / prod / v1 / files Try it Delete multiple files cURL ``` curl --request DELETE \ --url https://api.orion.file.ai/prod/v1/files \ --header 'Content-Type: application/json' \ --header 'x-api-key: ' \ --data ' { "fileIds": [ "8123612d-187c-44eb-8ac6-214e68506b7e", "8123612d-187c-44eb-8ac6-214e68506b7f" ] } ' ``` ``` import requests url = "https://api.orion.file.ai/prod/v1/files" payload = { "fileIds": ["8123612d-187c-44eb-8ac6-214e68506b7e", "8123612d-187c-44eb-8ac6-214e68506b7f"] } headers = { "x-api-key": "", "Content-Type": "application/json" } response = requests.delete(url, json=payload, headers=headers) print(response.text) ``` ``` const options = { method: 'DELETE', headers: {'x-api-key': '', 'Content-Type': 'application/json'}, body: JSON.stringify({ fileIds: ['8123612d-187c-44eb-8ac6-214e68506b7e', '8123612d-187c-44eb-8ac6-214e68506b7f'] }) }; fetch('https://api.orion.file.ai/prod/v1/files', options) .then(res => res.json()) .then(res => console.log(res)) .catch(err => console.error(err)); ``` ``` "https://api.orion.file.ai/prod/v1/files", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 30, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "DELETE", CURLOPT_POSTFIELDS => json_encode([ 'fileIds' => [ '8123612d-187c-44eb-8ac6-214e68506b7e', '8123612d-187c-44eb-8ac6-214e68506b7f' ] ]), CURLOPT_HTTPHEADER => [ "Content-Type: application/json", "x-api-key: " ], ]); $response = curl_exec($curl); $err = curl_error($curl); curl_close($curl); if ($err) { echo "cURL Error #:" . $err; } else { echo $response; } ``` ``` package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.orion.file.ai/prod/v1/files" payload := strings.NewReader("{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}") req, _ := http.NewRequest("DELETE", url, payload) req.Header.Add("x-api-key", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(string(body)) } ``` ``` HttpResponse response = Unirest.delete("https://api.orion.file.ai/prod/v1/files") .header("x-api-key", "") .header("Content-Type", "application/json") .body("{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}") .asString(); ``` ``` require 'uri' require 'net/http' url = URI("https://api.orion.file.ai/prod/v1/files") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Delete.new(url) request["x-api-key"] = '' request["Content-Type"] = 'application/json' request.body = "{\n \"fileIds\": [\n \"8123612d-187c-44eb-8ac6-214e68506b7e\",\n \"8123612d-187c-44eb-8ac6-214e68506b7f\"\n ]\n}" response = http.request(request) puts response.read_body ``` 200 401 403 404 422 429 500 ``` { "success": true } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "Invalid fileIds.", "instance": "/v1/{route}", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [], "message": "Invalid fileIds.", "error": "Unprocessable Entity", "statusCode": 422 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` ``` { "type": "https://errors.file.ai/validation-failed", "title": "Validation failed", "status": 422, "detail": "One or more fields are invalid.", "instance": "/v1/files/upload", "code": "VALIDATION_FAILED", "requestId": "0f1c8e03-978e-40d5-bc93-6894a57f9324", "errors": [ { "code": "REQUIRED", "detail": "must not be empty", "pointer": "#/fileName", "parameter": "limit" } ], "message": "Validation failed", "error": "Unprocessable Entity", "statusCode": 422, "retryAfter": 42 } ``` The **fileIds** field is used to specify the files that should be deleted. The files that are specified in the **fileIds** field will be deleted. If a files is not specified in the **fileIds** field, the files will not be deleted. ## [​ ](#overview) Overview Permanently deletes multiple files from the system. This is a bulk operation that allows you to remove several files in a single request. Warning: This action is irreversible. ## [​ ](#request-parameters) Request Parameters ### [​ ](#request-body) Request Body The request body must contain a JSON object with the following structure: Property Type Required Description fileIds string Yes Comma-separated list of file UUIDs to be deleted ## [​ ](#request-examples) Request Examples ### [​ ](#single-file-deletion) Single File Deletion ``` { "fileIds": "8123612d-187c-44eb-8ac6-214e68506b7e" } ``` ### [​ ](#multiple-files-deletion) Multiple Files Deletion ``` { "fileIds": "8123612d-187c-44eb-8ac6-214e68506b7e,8123612d-187c-44eb-8ac6-214e68506b7f,9234723e-298d-55fc-9bd7-325f79617c8g" } ``` ### [​ ](#success-response-200-ok) Success Response (200 OK) ``` { "success": true } ``` ### [​ ](#error-responses) Error Responses #### [​ ](#400-bad-request) 400 Bad Request ``` { "error": "Invalid request", "message": "fileIds parameter is required and cannot be empty" } ``` #### [​ ](#403-bad-request) 403 Bad Request ``` { "message": "Access denied. You are in readonly mode.", "error": "Forbidden", "statusCode": 403 } ``` #### [​ ](#422-bad-request) 422 Bad Request ``` { "message": "Invalid fileIds.", "error": "Unprocessable Entity", "statusCode": 422 } ``` ## [​ ](#important-notes) Important Notes * **Irreversible Action**: Deleted files cannot be recovered. Ensure you have proper backups before deletion. * **UUID Format**: File IDs must be valid UUIDs in the format xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx. * **Comma Separation**: When providing multiple file IDs as a string, separate them with commas without spaces. * **Batch Processing**: The endpoint processes all file IDs in a single transaction where possible. * **Partial Failures**: If some files cannot be deleted, the operation will complete for valid files and return error details for failed ones. * **Permission Checks**: Each file deletion is subject to individual permission validation. ## [​ ](#file-id-requirements) File ID Requirements * Must be valid UUID format (36 characters including hyphens) * File must exist in the system * User must have delete permissions for each file * Maximum recommended batch size: 100 files per request ## [​ ](#use-cases) Use Cases * Bulk cleanup of temporary or processed files * User-initiated deletion of selected files from file manager * Automated cleanup processes removing expired files * Storage management operations to free up space * Data privacy compliance for user data deletion requests ## [​ ](#best-practices) Best Practices * Validate file ownership before deletion * Implement confirmation dialogs for user-facing operations * Log deletion activities for audit purposes * Check file dependencies before deletion * Use reasonable batch sizes to avoid timeouts * Handle partial failures gracefully in your application logic This endpoint accepts an optional `Idempotency-Key` request header so a retry cannot apply the change twice. See [Idempotent requests](/docs-api/api-idempotency). #### Authorizations [​ ](#authorization-x-api-key) x-api-key string header required API key for authentication #### Headers [​ ](#parameter-idempotency-key) Idempotency-Key string Optional opaque key (max 255 chars, UUIDv4 recommended) making this request idempotent for 24h: a retry with the same key and body replays the original response with Idempotent-Replayed: true. #### Body application/json Delete multiple files input The body is of type `string`. #### Response 200 application/json The files are deleted [​ ](#response-success) success boolean [ Update file export status ](/api-reference/endpoint/update-file-export-status)[ Upload a file ](/api-reference/endpoint/upload-a-file) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # Upload a file Source: https://docs.file.ai/api-reference/endpoint/upload-a-file post /prod/v1/files/upload Upload a file to the server, this will return a presigned upload url to be used for the upload. The presignedUploadURL is valid for 300 seconds (5 minutes) and can be used multiple times. The input should contain the following information: - `fileName`: the name of the file to be uploaded - `fileType`: the type of the file to be uploaded - `isSplit`: whether the file is a split file or not (optional, default: false) - `isSplitExcel`: whether to split Excel files by worksheets (optional, default: false) - `callbackURL`: (optional) the URL that will be called after file processing. Must be a valid HTTPS URL. - If provided, this URL takes precedence over the API key's default callback URL - If not provided, the API key's default callback URL will be used (if configured) - The response includes `callbackURLSource` indicating whether the URL came from the request ("user") or API key ("api_key") - `ocrModel`: the OCR model to be used for file processing (optional). Available models: - **English Models:** - `Beethoven_ENG_O5.6` - OpenAI v6 - `Beethoven_ENG_G5.5` - Gemini v5 - `Beethoven_ENG_GP25` - Gemini Pro 2.5 - `Beethoven_ENG_GP25.1` - Gemini Pro 2.5 v1 - `Beethoven_ENG_GP25.2` - Gemini Pro 2.5 PDF - `Beethoven_CUS_O5.1` - Custom OpenAI v8 - `Beethoven_CUS_O5.2` - Custom Gemini v13 - `Unified (google-document-ai-ocr-gemini-v10)` - Unified model - `Aegis (google-document-ai-ocr-gemini-aegis-v1)` - Aegis model - **Chinese Models:** - `Beethoven_ZH_O5.9` - Chinese OpenAI v9 - **Japanese Models:** - `Beethoven_JP_O5.3` - Japanese OpenAI v3 - `Beethoven_JP_G5.4` - Japanese Gemini fine-tuned - **Thai Models:** - `Beethoven_TH_O5.1` - Thai OpenAI v1 - `Beethoven_TH_G5.1` - Thai Gemini v1 - `schemaLocking`: whether the schema should be locked after the file is uploaded, must be one of true or false (optional) - `directoryId`: the directory id where the file should be uploaded (optional) - `destinationPath`: slash-delimited folder path where the file should be placed (e.g. "mammals/walrus"). Folders are auto-created if they do not exist. Can be used together with directoryId (optional) - `isEphemeral`: whether the file and all related data should be deleted after the file is processed, must be one of true or false (optional, default: false) - `pageCount`: page count of the file, used for early validation against page limits (optional) --- # Upload a large file using multipart upload Source: https://docs.file.ai/api-reference/endpoint/upload-a-large-file-using-multipart-upload post /prod/v1/files/upload/multipart Initiate a multipart upload for large files (typically >100MB). This will return presigned URLs for each part. Each presigned URL is valid for 900 seconds (15 minutes) and can be used multiple times. The workflow is: 1. Call this endpoint to get presigned URLs for each part 2. Upload each part to its respective presigned URL using PUT requests 3. Call the complete multipart upload endpoint with all part ETags The input should contain: - `fileName`: the name of the file to be uploaded - `fileType`: the MIME type of the file - `fileSize`: the size of the file in MB - `partSizeLimit`: (optional) the size limit for each part in MB - `isSplit`: whether the file should be split after upload (optional, default: false) - `isSplitExcel`: whether to split Excel files by worksheets (optional, default: false) - `callbackURL`: the url that will be called after processing (optional) - `ocrModel`: the OCR model to use (optional) - `schemaLocking`: whether the schema should be locked (optional) - `directoryId`: the directory id where the file should be uploaded (optional) - `destinationPath`: slash-delimited folder path where the file should be placed (e.g. "mammals/walrus"). Folders are auto-created if they do not exist. Can be used together with directoryId (optional) - `isEphemeral`: whether the file and all related data should be deleted after the file is processed, must be one of true or false (optional, default: false) - `pageCount`: page count of the file, used for early validation against page limits (optional) --- # Complete a multipart upload Source: https://docs.file.ai/api-reference/endpoint/complete-a-multipart-upload post /prod/v1/files/upload/multipart/complete Complete a multipart upload after all parts have been uploaded to S3. After uploading all parts using the presigned URLs from the initiate endpoint, call this endpoint with the ID, key, and all part ETags to finalize the upload. The input should contain: - `id`: the ID from the multipart upload initiation (previously called uploadId) - `key`: the S3 key from the multipart upload initiation - `parts`: array of objects with partNumber and eTag for each uploaded part --- # Get a file upload processing steps for that uploadId Source: https://docs.file.ai/api-reference/endpoint/get-a-file-upload-processing-steps-for-that-uploadid get /prod/v1/uploads/{uploadId} Get a file upload processing steps for that uploadId --- # Create a file preview Source: https://docs.file.ai/api-reference/endpoint/create-a-file-preview post /prod/v1/files/preview Create a file preview using a S3 Path. The input should contain the following information: - `s3Path`: the path of the file in the S3 bucket --- # Create file previews in batch Source: https://docs.file.ai/api-reference/endpoint/create-file-previews-in-batch post /prod/v1/files/preview/batch Create presigned preview URLs for several S3 paths in a single request. Accepts up to 100 paths in `s3Paths`. Results are returned in request order, each echoing its `s3Path` so callers can map URLs back to files. All-or-nothing: if any path is invalid the whole request fails — no partial results. Use `POST /v1/files/preview` for a single path. --- # Update a file's schema field definitions Source: https://docs.file.ai/api-reference/endpoint/update-file-schema patch /prod/v1/files/schema This endpoint is used to patch the schema fields that are associated with the file types in the workspace. The response will contain a list of schemas that are associated with the file types in the workspace. The `fileId` field is used to specify the file that should be changed to the new schema version. The `schemaId` field is used to specify the schema that should be updated. The `shouldCreateCustomIfNotExists` field is used to specify whether a custom schema should be created if it does not exist. The `schemaVersionType` field is used to specify the type of schema version to be create. The allowed values are `MAJOR`, `MINOR`, and `PATCH`. The `shouldChangeDocumentToNewSchemaVersion` field is used to specify whether the document should be changed to the new schema version. The `fieldPaths` field is used to specify the fields that should be updated. --- # Get file schema values by fileIds Source: https://docs.file.ai/api-reference/endpoint/get-file-schema-values-by-fileids get /prod/v1/files/{fileId}/values Get file schema values by fileIds of the current organization and workspace. This endpoint is paginated, you can use the `page` and `limit` query parameters to control the pagination. For table/list fields with more than 100 rows, use `includeFullTableData=true` to retrieve all rows from S3 storage. --- # Update a file's extracted values (fields and table/list cells) Source: https://docs.file.ai/api-reference/endpoint/update-file-values patch /prod/v1/files/{fileId}/values Update document form values for a file, for both scalar fields and table/list cells. The `schemaValueId` field is the `id` returned by `GET /files/:fileId/values` (the form value id). The `fieldPaths` field is an array of updates. Each entry must specify **exactly one** of: - `value`: the new value for a scalar field. The `path` is the field title (the leading slash is optional, e.g. `/Supplier country`). - `cells`: cell-level updates for a table/list field. The `path` is the table field title (e.g. `Line_items`). Each cell contains: - `rowIndex`: the zero-based row index within the table. - `column`: the column title (must match one of the table columns). - `value`: the new cell value. Table updates are a full read-modify-write: the complete CSV-backed rows are loaded, the specified cells are applied, and the table is re-persisted (updating the preview, document history, audit trail, and re-flattening the document). --- # Get file ocr Source: https://docs.file.ai/api-reference/endpoint/get-file-ocr get /prod/v1/files/{fileId}/ocr Get file ocr (Optical Character Recognition). The `fileId` field is used to specify the file ocr that should be retrieved. --- # Rerun file form filling Source: https://docs.file.ai/api-reference/endpoint/rerun-file-form-filling patch /prod/v1/files/schema/rerun This endpoint is used to rerun the form filling of a file. It is used to update the file fields based on the current file content and the current schema version. The `fileId` field is used to specify the file that should be updated. The `backgroundRerun` field is used to determine whether the form filling should be run in the background. If set to `true`, the form filling will be run in the background. If set to `false`, the form filling will be run synchronously. The `rerunOptions` field is used to specify the options for the form filling. The allowed fields are `paths`, `partialRerun`, and `notification`. The `paths` field is used to specify the fields that should be updated. The fields that are specified in the `paths` field will be updated with the new schema fields. The `partialRerun` field is used to determine whether the form filling should only be run for the modified fields. If set to `true`, the form filling will only be run for the modified fields. If set to `false`, the form filling will be run for all fields. The `notification` field is used to specify the notification options for the form filling. The allowed fields are `shouldSendOnCompletion`, `referenceEntityName`, `referenceEntityId`, and `referenceEntityAction`. The `shouldSendOnCompletion` field is used to determine whether a notification should be sent when the form filling is completed. If set to `true`, a notification will be sent when the form filling is completed. If set to `false`, no notification will be sent. The `referenceEntityName` field is used to specify the name of the reference entity that should be used for the notification. The allowed values are `FILE` and `SCHEMA`. The `referenceEntityId` field is used to specify the id of the reference entity that should be used for the notification. The `referenceEntityAction` field is used to specify the action that should be used for the notification. The allowed values are `CREATE`, `UPDATE`, and `DELETE`. --- # Get file types Source: https://docs.file.ai/api-reference/endpoint/get-file-types get /prod/v1/file-type This endpoint is used to get all the file types that are available in the workspace. The response will contain a list of file types that are associated with the workspace and their respective properties. --- # Approve a file type Source: https://docs.file.ai/api-reference/endpoint/approve-a-file-type patch /prod/v1/file-type/approve This endpoint is used to approve a file type. The response will contain the success status. --- # Rename a file type Source: https://docs.file.ai/api-reference/endpoint/rename-a-file-type patch /prod/v1/file-type/rename This endpoint is used to rename a file type. The response will contain the success status. --- # Get all schemas Source: https://docs.file.ai/api-reference/endpoint/get-all-schemas Get all schemas cURL ``` curl --request GET \ --url https://api.orion.file.ai/prod/v1/schemas \ --header 'x-api-key: ' ``` ``` import requests url = "https://api.orion.file.ai/prod/v1/schemas" headers = {"x-api-key": ""} response = requests.get(url, headers=headers) print(response.text) ``` ``` const options = {method: 'GET', headers: {'x-api-key': ''}}; fetch('https://api.orion.file.ai/prod/v1/schemas', options) .then(res => res.json()) .then(res => console.log(res)) .catch(err => console.error(err)); ``` ``` "https://api.orion.file.ai/prod/v1/schemas", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 30, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "GET", CURLOPT_HTTPHEADER => [ "x-api-key: " ], ]); $response = curl_exec($curl); $err = curl_error($curl); curl_close($curl); if ($err) { echo "cURL Error #:" . $err; } else { echo $response; } ``` ``` package main import ( "fmt" "net/http" "io" ) func main() { url := "https://api.orion.file.ai/prod/v1/schemas" req, _ := http.NewRequest("GET", url, nil) req.Header.Add("x-api-key", "") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(string(body)) } ``` ``` HttpResponse response = Unirest.get("https://api.orion.file.ai/prod/v1/schemas") .header("x-api-key", "") .asString(); ``` ``` require 'uri' require 'net/http' url = URI("https://api.orion.file.ai/prod/v1/schemas") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Get.new(url) request["x-api-key"] = '' response = http.request(request) puts response.read_body ``` 200 offset ``` { "schemas": [ { "id": "5f5f726eea75272d54e1e1e1", "fields": [ { "id": "5f5f726eea75272d54e1e1e2" } ], "blueprintId": "5f5f726eea75272d54e1e1e3", "originalSchemaId": "5f5f726eea75272d54e1e1e4", "schemaHistoryId": "5f5f726eea75272d54e1e1e5", "schemaHistoryVersionNumber": 1, "fileId": "5f5f726eea75272d54e1e1e6", "blueprintMode": "USER", "schemaType": "SYSTEM", "state": "ACTIVE", "usageApprovalState": "APPROVED", "schemaOverrideType": "OVERRIDE", "aiSchemaVersion": "5f5f726eea75272d54e1e1e7", "createdAt": "2020-10-05T14:30:00.000Z", "updatedAt": "2020-10-05T14:30:00.000Z" } ], "count": 1, "currentPage": 1 } ``` Endpoints # Get all schemas Copy pageCopy page This endpoint is used to get all schemas. The response will contain a list of all schemas that are available for the current user in the current workspace. Pass the `cursor` query parameter (empty for the first page) to use cursor pagination: the response envelope becomes `{ data, pagination: { limit, nextCursor, hasMore } }`. Follow `nextCursor` until `hasMore` is false; keep the same sort and filter parameters for every page of a walk. The `schemas` field is an array of schemas. Each schema contains the following fields: * `id`: The ID of the schema. This is the unique identifier for the schema. * `fields`: An array of fields. Each field contains the following fields: * `id`: The ID of the field. This is the unique identifier for the field. * `blueprintId`: The ID of the blueprint that the schema is associated with. * `originalSchemaId`: The ID of the original schema that the schema is based on. * `schemaHistoryId`: The ID of the schema history that the schema is associated with. * `schemaHistoryVersionNumber`: The version number of the schema history that the schema is associated with. This is a number that increments every time a new version of the schema is created. * `fileId`: The ID of the file that the schema is associated with. * `blueprintMode`: The mode of the blueprint that the schema is associated with. This can be either USER or SYSTEM. * `schemaType`: The type of the schema. This can be either SYSTEM, CUSTOM, or AI. * `state`: The state of the schema. This can be either ACTIVE or INACTIVE. * `usageApprovalState`: The state of the usage approval of the schema. This can be either APPROVED, PENDING, or REJECTED. * `schemaOverrideType`: The type of the schema override. This can be either OVERRIDE or DEFAULT. * `aiSchemaVersion`: The version of the AI schema that the schema is associated with. * `createdAt`: The date and time when the schema was created. * `updatedAt`: The date and time when the schema was last updated. * `deletedAt`: The date and time when the schema was deleted. * `deletedBy`: The ID of the user who deleted the schema. Copy pageCopy page GET https://api.orion.file.aihttps://api.orion.{instance}.file.ai / prod / v1 / schemas Try it Get all schemas cURL ``` curl --request GET \ --url https://api.orion.file.ai/prod/v1/schemas \ --header 'x-api-key: ' ``` ``` import requests url = "https://api.orion.file.ai/prod/v1/schemas" headers = {"x-api-key": ""} response = requests.get(url, headers=headers) print(response.text) ``` ``` const options = {method: 'GET', headers: {'x-api-key': ''}}; fetch('https://api.orion.file.ai/prod/v1/schemas', options) .then(res => res.json()) .then(res => console.log(res)) .catch(err => console.error(err)); ``` ``` "https://api.orion.file.ai/prod/v1/schemas", CURLOPT_RETURNTRANSFER => true, CURLOPT_ENCODING => "", CURLOPT_MAXREDIRS => 10, CURLOPT_TIMEOUT => 30, CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1, CURLOPT_CUSTOMREQUEST => "GET", CURLOPT_HTTPHEADER => [ "x-api-key: " ], ]); $response = curl_exec($curl); $err = curl_error($curl); curl_close($curl); if ($err) { echo "cURL Error #:" . $err; } else { echo $response; } ``` ``` package main import ( "fmt" "net/http" "io" ) func main() { url := "https://api.orion.file.ai/prod/v1/schemas" req, _ := http.NewRequest("GET", url, nil) req.Header.Add("x-api-key", "") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(string(body)) } ``` ``` HttpResponse response = Unirest.get("https://api.orion.file.ai/prod/v1/schemas") .header("x-api-key", "") .asString(); ``` ``` require 'uri' require 'net/http' url = URI("https://api.orion.file.ai/prod/v1/schemas") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Get.new(url) request["x-api-key"] = '' response = http.request(request) puts response.read_body ``` 200 offset ``` { "schemas": [ { "id": "5f5f726eea75272d54e1e1e1", "fields": [ { "id": "5f5f726eea75272d54e1e1e2" } ], "blueprintId": "5f5f726eea75272d54e1e1e3", "originalSchemaId": "5f5f726eea75272d54e1e1e4", "schemaHistoryId": "5f5f726eea75272d54e1e1e5", "schemaHistoryVersionNumber": 1, "fileId": "5f5f726eea75272d54e1e1e6", "blueprintMode": "USER", "schemaType": "SYSTEM", "state": "ACTIVE", "usageApprovalState": "APPROVED", "schemaOverrideType": "OVERRIDE", "aiSchemaVersion": "5f5f726eea75272d54e1e1e7", "createdAt": "2020-10-05T14:30:00.000Z", "updatedAt": "2020-10-05T14:30:00.000Z" } ], "count": 1, "currentPage": 1 } ``` The response will contain a list of all schemas that are available for the current user in the current workspace. The schemas field is an array of schemas. Each **schema** contains the following fields: * **id**: The ID of the schema. This is the unique identifier for the schema. * **fields**: An array of fields. Each **field** contains the following fields: * **id**: The ID of the field. This is the unique identifier for the field. * **blueprintId**: The ID of the blueprint that the schema is associated with. * **originalSchemaId**: The ID of the original schema that the schema is based on. * **schemaHistoryId**: The ID of the schema history that the schema is associated with. * **schemaHistoryVersionNumber**: The version number of the schema history that the schema is associated with. This is a number that increments every time a new version of the schema is created. * **fileId**: The ID of the file that the schema is associated with. * **blueprintMode**: The mode of the blueprint that the schema is associated with. This can be either USER or SYSTEM. * **schemaType**: The type of the schema. This can be either SYSTEM, CUSTOM, or AI. * **state**: The state of the schema. This can be either ACTIVE or INACTIVE. * **usageApprovalState**: The state of the usage approval of the schema. This can be either APPROVED, PENDING, or REJECTED. * **schemaOverrideType**: The type of the schema override. This can be either OVERRIDE or DEFAULT. * **aiSchemaVersion**: The version of the AI schema that the schema is associated with. ## [​ ](#query-parameters) Query Parameters Parameter Type Required Default Description schemaIds string No — Comma-separated schema ids to fetch uploadIds string No — Comma-separated upload ids to fetch schemas for page number No 1 Page number for pagination limit number No 100 Number of items per page sortBy string No `createdAt` Field to sort by sortOrder string No `ASC` Sort direction (`ASC` or `DESC`) cursor string No — Opaque cursor. Send it (empty for the first page) to switch to cursor pagination. This endpoint supports both pagination modes. Send the `cursor` parameter (empty value for the first page) to use cursor pagination; omit it entirely to keep the legacy `page`/`limit` response shape. See [Pagination](/docs-api/api-pagination) for the full comparison. In cursor mode the response becomes `{ data, pagination: { limit, nextCursor, hasMore } }`. Follow `nextCursor` until `hasMore` is `false`, keeping the same sort and filter parameters for every page of the walk. Omit `cursor` and the response keeps its `{ schemas, count, currentPage }` shape. #### Authorizations [​ ](#authorization-x-api-key) x-api-key string header required API key for authentication #### Query Parameters [​ ](#parameter-schema-ids) schemaIds string Schema Ids Example: `"649e2d2d2d2d2d2d2d2d2d2d"` [​ ](#parameter-upload-ids) uploadIds string Upload Ids Example: `"649e2d2d2d2d2d2d2d2d2d2d"` [​ ](#parameter-page) page number default:1 Page number Example: `1` [​ ](#parameter-limit) limit number default:100 Number of items per page Example: `100` [​ ](#parameter-sort-by) sortBy string default:createdAt Field to sort by Example: `"createdAt"` [​ ](#parameter-sort-order) sortOrder enum default:ASC Sort direction Available options: `ASC`, `DESC` Example: `"ASC"` [​ ](#parameter-cursor) cursor string Opaque pagination cursor. Provide the parameter (empty value for the first page) to switch to cursor pagination: the response becomes { data, pagination: { limit, nextCursor, hasMore } } with default limit 50 (max 100). Omit it entirely for legacy page/limit pagination. In cursor mode sortBy must be one of createdAt, updatedAt, \_id, and cursors are bound to the sortBy/sortOrder they were minted with. #### Response 200 application/json The list of schemas. Offset mode (default) returns `{schemas, count, currentPage}`; cursor mode (when the `cursor` query parameter is present) returns `{data, pagination}`. * Option 1 * Option 2 [​ ](#response-one-of-0-schemas) schemas object\[\] required [​ ](#response-one-of-0-count) count number required Total number of matching schemas [​ ](#response-one-of-0-current-page) currentPage number required [ Rename a file type ](/api-reference/endpoint/rename-a-file-type)[ Patch a schema ](/api-reference/endpoint/patch-a-schema) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # [Deprecated] Patch workspace schema field definitions Source: https://docs.file.ai/api-reference/endpoint/patch-a-schema patch /prod/v1/schemas **Deprecated.** To update extracted field values of a file — including regular fields (string, number, enum, date, boolean) and list/table cell values — use `PATCH /files/{fileId}/values` instead. This endpoint remains available for patching schema field definitions, but will be removed in a future version. This endpoint is used to patch a schema. The `schemaId` field is used to specify the schema that should be patched. The `fields` field is used to specify the fields that should be updated. The `fields` field is an array of fields. Each field contains the following fields: - `title`: The title of the field. This is the human-readable name of the field. - `description`: The description of the field. This is a human-readable description of the field. - `type`: The type of the field. This can be either STRING, NUMBER, BOOLEAN, DATE, or OBJECT. - `format`: The format of the field. This is a string that indicates the format of the field. - `hidden`: Whether the field is hidden or not. This is a boolean value that indicates whether the field is hidden or not. - `internet_context`: Whether the field is internet context or not. This is a boolean value that indicates whether the field is internet context or not. - `fieldSuggestion`: The field suggestion of the field. This is an object that contains the following fields: - `id`: The ID of the field suggestion. This is the unique identifier for the field suggestion. - `name`: The name of the field suggestion. This is the human-readable name of the field suggestion. - `items`: The items of the field. This is an array of fields that are the items of the field. - `value`: The value of the field. This is the value of the field. - `enumValues`: The enum values of the field. This is an array of strings that are the enum values of the field. - `inlineListType`: The inline list type of the field. This is a string that indicates the inline list type of the field. --- # Get file type schemas Source: https://docs.file.ai/api-reference/endpoint/get-file-type-schemas get /prod/v1/schemas/file-type This endpoint is used to get the file type schemas that are available in the workspace. The response will contain a list of schemas that are associated with the file types in the workspace. --- # Get project setting Source: https://docs.file.ai/api-reference/endpoint/get-project-setting get /prod/v1/project/setting This endpoint is used to get the setting of the project. The response will contain the setting of the project. --- # Update project setting Source: https://docs.file.ai/api-reference/endpoint/update-project-setting patch /prod/v1/project/setting This endpoint is used to update the setting of the project. The response will contain the updated setting of the project. --- # Get all directories (folders) Source: https://docs.file.ai/api-reference/endpoint/get-all-directories get /prod/v1/directories Retrieve all directories (folders) in the workspace with hierarchical structure. This endpoint returns a paginated list of directories including their nested subfolders. Each directory includes breadcrumb navigation for easy hierarchy tracking. Pagination is over TOP-LEVEL directories: one page row is one root folder, returned with its complete subtree. A subfolder is therefore never a page row of its own — reach it through its root's `subfolders`, or use cursor mode for a flat walk of every directory. Because a page carries whole subtrees, page size in bytes is driven by how deep those trees are; prefer cursor mode for large workspaces. Query Parameters: - page (optional): Page number for pagination (default: 1) - limit (optional): Number of TOP-LEVEL directories per page (default: 100). Subfolders nested under a returned root do not count against it, so this bounds rows, NOT total directories returned — a page whose subtrees expand past 5000 directories is rejected with 400; lower `limit` or use cursor mode. - search (optional): Case-insensitive substring match on the directory name. When present, `directories` is a FLAT list of matches — no `subfolders` nesting — and each match carries its full `breadcrumbs` path. Omit it to get the nested tree. Response: The response includes: - directories: Array of directory objects with nested subfolders - count: Total number of TOP-LEVEL directories, i.e. the number of page rows across all pages. Divide by `limit` for the page count. - currentPage: Current page number Each directory object contains: - Basic information: id, name, description, status - Hierarchy information: parentId, depth, breadcrumbs, subfolders - Metadata: userId, organizationId, workspaceId - Timestamps: createdAt, updatedAt - isFromAutomatedRule: Indicates if directory was created by automation Notes: - Directories are returned with their complete subfolder hierarchy - Breadcrumbs provide the full path from root to current directory - Subfolders array contains nested directories recursively Cursor pagination: pass the `cursor` query parameter (empty for the first page) and the response becomes `{ data, pagination: { limit, nextCursor, hasMore } }` where `data` is a FLAT list of directories — no `subfolders` nesting; each row carries its `breadcrumbs` ancestor path instead. Follow `nextCursor` until `hasMore` is false. Cursors are bound to the `search` value they were minted with; keep it identical for every page of a walk. --- # Create a new directory (folder) Source: https://docs.file.ai/api-reference/endpoint/create-a-directory post /prod/v1/directories Create a new directory (folder) to organize your files within the workspace. Directories can be nested by providing a parentId to create subdirectories. The system enforces a maximum directory depth limit (configurable via system settings, default is 1 level). Request Body Parameters: - name (required): The name of the directory - description (optional): A description for the directory - parentId (optional): The ID of the parent directory to create a subdirectory. If not provided, the directory will be created at the root level Response: The response includes the created directory object with: - Basic information: id, name, description, status - Hierarchy information: parentId, depth - Metadata: userId, organizationId, workspaceId, createdBy, lastModifiedBy, etc. - Timestamps: createdAt, updatedAt Notes: - The directory is created with status: "active" by default - The authenticated user becomes the owner of the directory - Directory depth is automatically calculated based on parent hierarchy --- # Get workspace subscription Source: https://docs.file.ai/api-reference/endpoint/get-workspace-subscription get /prod/v1/subscription This endpoint is used to get the subscription of the workspace. The response will contain the subscription of the workspace. --- # Export schemas Source: https://docs.file.ai/api-reference/endpoint/export-schemas post /prod/v1/files/schema/export Export one or more schemas to a downloadable zip file. **Request Body:** - `aiSchemaIds` (string, required): Comma-separated list of schema IDs to export **Response:** - `exportedFileDownloadUrl` (string): Pre-signed URL to download the exported zip file (valid for 1 hour) - `s3Path` (string): S3 path where the file is stored - `fileName` (string): Name of the exported file **Example:** ```json { "aiSchemaIds": "649e2d2d2d2d2d2d2d2d2d2d,649e2d2d2d2d2d2d2d2d2d2e" } ``` --- # Import schemas Source: https://docs.file.ai/api-reference/endpoint/import-schemas post /prod/v1/files/schema/import Import schemas from a previously exported zip file. **Request:** - Content-Type: `multipart/form-data` - `file` (file, required): The exported zip file **Response:** - `success` (boolean): Whether the import was successful - `importedBlueprintsCount` (number): Number of blueprints imported - `importedSchemasCount` (number): Number of schemas imported - `conflictBlueprintsCount` (number): Number of blueprints that had conflicts - `conflictBlueprints` (string[]): List of conflicting blueprint class names - `blueprintIds` (string[]): IDs of the imported blueprints - `schemaIds` (string[]): IDs of the imported schemas **Notes:** - All imported blueprints and schemas will be assigned new IDs - Metadata (organization, workspace, user) will be updated to the importing user - If a blueprint with the same class name already exists, the import will fail --- # Execute a query builder configuration Source: https://docs.file.ai/api-reference/endpoint/execute-a-query post /prod/v1/query/execute Executes a query builder configuration using the workspace and organization bound to the API key. --- # Get workflow by id Source: https://docs.file.ai/api-reference/endpoint/get-workflow-by-id get /prod/v1/workflows/{workflowId} Retrieve a single workflow by its ID within the current organization and workspace. The response includes workflow metadata and a summary of each workflow step (id, type, execution order, and level). Internal step configuration is not exposed. --- # Get workflow result groups by unique identifiers Source: https://docs.file.ai/api-reference/endpoint/get-workflow-result-groups get /prod/v1/workflows/result-groups Retrieve workflow execution result groups for one or more execution group unique identifiers (e.g. folder or subfolder ids). Query parameters: - executionGroupUniqueIdentifiers (required): Comma-separated list of unique identifiers - workflowStepIds (optional): Comma-separated workflow step ObjectIds to filter results Each result group includes structural fields and, when present, an assessment summary with overallPassed, assessmentStatus, rejectionReason, and per-rule pass/fail. --- # fileAI MCP Server overview Source: https://docs.file.ai/docs-mcp/get-started-mcp ## On this page * [Features](#features) * [Requirements](#requirements) fileAI MCP # fileAI MCP Server overview Copy pageCopy page What can our MCP server do? Copy pageCopy page The **fileAI MCP Server** offers a robust set of tools to work with the fileAI file processing pipeline. It allows for uploading files, performing **Optical Character Recognition (OCR)**, classifying documents, and extracting **structured data**. The server leverages the **Model Context Protocol (MCP)** to provide seamless integration with AI models, enabling them to work with your documents **programmatically**. ## [​ ](#features) **Features** ## End-to-end file processing * From file upload to structured data extraction — manage the entire lifecycle of your files. ## AI-powered * Leverage powerful AI models for OCR, file classification, and structured data extraction. ## Schema management * Define, update, and manage schemas to control how data is extracted for your specific needs. ## Asynchronous processing * Upload files and track processing status asynchronously. ## [​ ](#requirements) **Requirements** To use the fileAI MCP Server, make sure you have the following items at hand: ## A fileAI Account You need an account to use our API endpoints and MCP servers. Sign up or login [here](orion.file.ai). ## A fileAI API key * After creating your fileAI account, you can generate your API Key. ## Verified AI Schemas fileAI suggests extraction and data fetch schemas. Confirm or edit these in the UI to call them directly via MCP. ## Your MCP toolkit > An MCP client (such as Cursor or Claude Desktop) and **npx** (included with Node.js) available in your environment. [ Set Up and Configuration ](/docs-mcp/mcp-server-configuration) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined) --- # MCP Server Configuration Source: https://docs.file.ai/docs-mcp/mcp-server-configuration ## On this page * [Configuring Cursor](#configuring-cursor) * [Configuring Claude](#configuring-claude) * [Configuring Windsurf, Cline Github Copilot and other MCP clients](#configuring-windsurf-cline-github-copilot-and-other-mcp-clients) fileAI MCP # MCP Server Configuration Copy pageCopy page How to run our MCP server on Cursor and Claude. Copy pageCopy page ## [​ ](#configuring-cursor) **Configuring Cursor** **HTTP Transport** 1. Open Cursor Settings (Ctrl/Cmd + ,) 2. Navigate to Tools & Integration → MCP Tools 3. Click **\+ New MCP Server** ![Screenshot 2025-08-18 at 9.55.22 AM.png](/images/Screenshot2025-08-18at9.55.22AM.png) **Use the configuration below:** ``` { "mcpServers": { "FileAI-Prod": { "url": "https://mcp.file.ai", "env": { "API_KEY": "YOUR_API_KEY_HERE" } } } } ``` Your mcp.json should look like below: ![Screenshot 2025-08-18 at 9.56.25 AM.png](/images/Screenshot2025-08-18at9.56.25AM.png) Verify that the MCP Server is set up correctly. You should see that tools are enabled. ![Screenshot 2025-08-18 at 11.32.45 AM.png](/images/Screenshot2025-08-18at11.32.45AM.png) ## [​ ](#configuring-claude) **Configuring Claude** To add the fileAI MCP Server to Claude: 1. Go to your **Claude Integrations** page 2. Click **\+ Add integration** 3. In the URL field, enter: ``` https://mcp.file.ai/sse?x-api-key=Bearer YOUR_API_KEY_HERE ``` Replace `YOUR_API_KEY_HERE` with your actual fileAI API key. \_Please note that Claude has restrictive rules regarding the ability to send PII via MCP. Please ensure that your use case is consistent with Claude’s Terms of Service. \_ * * * ## [​ ](#configuring-windsurf-cline-github-copilot-and-other-mcp-clients) Configuring Windsurf, Cline Github Copilot and other MCP clients Visit [mcp.file.ai](http://mcp.file.ai) to get instructions on how to install fileAI’s MCP Server on other clients. [ Introduction ](/docs-mcp/get-started-mcp) [linkedin](https://www.linkedin.com/company/file-ai)[x](https://x.com/fileAI_)[youtube](https://www.youtube.com/@file_AI) [Powered byThis documentation is built and hosted on Mintlify, a developer documentation platform](https://www.mintlify.com?utm_campaign=poweredBy&utm_medium=referral&utm_source=undefined)