Workspaces lakehouse
The lakehouse collects usage information, document audit trails, Logic Pilot traces, and feedback from every workspace. This page shows where that data is saved, and which program reads it.
Follow a batch of data from the workspaces to the place that keeps it. Choose a type of data. The glowing cubes show the route.
About 7 seconds after a burst of changes, each workspace sends them to the worker.
Document audit
Document changes are sent when events happen. About 7 seconds after a burst of changes, the workspace sends them to the Cloudflare worker. A poll every 5 minutes is a safety net, in case an event is missed. There is no D1 for this save.
The worker sends the batch into Cloudflare Pipelines. About a minute later the rows land in Iceberg document_audit.v1, in the R2 bucket workspace-events. That send also clears the KV cache for the document.
Saved in the lakehouse
Terms used on this page
Programs and places
| Term | Meaning |
|---|---|
| Workspace | One customer system. It has its own database. |
| Workspace backend | The program that sends data from each workspace. See logic-bee. |
| Worker | The program that receives the data and saves it. See bob-workspace-orchestration. |
| Builder | The program that reads the lakehouse data for charts. See bob-api. |
| D1 | The hot store. A fast database. It answers queries about recent data. |
| Iceberg | The cold store. Files kept in the R2 bucket workspace-events. It keeps the long history. |
| R2 | Cloudflare's file storage. It holds the Iceberg files and the workspace file buckets. |
| Cloudflare Pipelines | The Cloudflare service that writes the Iceberg files. Rows are ready about 1 minute after the worker sends them. |
Lakehouse concepts
| Term | Meaning |
|---|---|
| Hot store | A store built for fast, recent queries. On this page, that is D1. |
| Cold store | A store built to hold the long history. On this page, that is Iceberg. |
| Cleanup sweep | The job that runs every 6 hours and deletes D1 rows older than 28 days. Iceberg keeps its own copy. |
| Deployment | Another name for one workspace. Its id is the workspace id. |
| Company | One company inside a workspace. A workspace can hold more than one. |
| Business unit | A group inside a company. |
Logic Pilot terms
| Term | Meaning |
|---|---|
| Logic Pilot | The platform's AI agent. People talk with it in conversations, and each run can be traced as spans. |
| Conversation | One Logic Pilot chat thread. It has a mode, a model, and a message count. |
| Session | The document that holds a person's ratings on Logic Pilot messages, identified by sessionDocId. |
| Trace | One Logic Pilot run, from start to end. |
| Span | One step inside a trace, for example one model call. |
| Root span | The first span in a trace. |
Other terms
| Term | Meaning |
|---|---|
| Burst | A short period of rapid changes to one document. |
| Bucket | One Cloudflare R2 container for files. Each workspace has a public (cdn) bucket and a private (storage) bucket. |
| Class A / Class B operations | Cloudflare's two cost categories for R2 operations. Class A covers writes and lists. Class B covers reads. |
| JSON Patch | The format used in jsonPatch to describe one document change, as a list of operations. |
The three programs
Three programs move this data.
| Program | What it does | Repository |
|---|---|---|
| Workspace backend | Sends data from each workspace | logic-bee |
| Worker | Receives the data and saves it | bob-workspace-orchestration |
| Builder backend | Reads the data for charts | bob-api |
A workspace is one customer system. It has its own database. There are many workspaces.
The workspace id in this store is the deployment id.
Local workspaces
A workspace on a developer machine does not push lakehouse data. NODE_ENV is not production, so the workspace returns before it calls the worker.
The scheduled sends still start. They stop before they read companies or open a request. A document change, a document count, a Logic Pilot span, a stats snapshot, a bug report, a feedback note, and a Logic Pilot rating stay on the local machine. They are not copied to D1 or Iceberg. An audit-trail read on that machine does not call the worker.
A production workspace does push this data. NODE_ENV is production.
File storage is separate. The worker collects those totals. The workspace does not send them in either environment.
What the lakehouse is
The lakehouse is one shared store. The business documents stay in the workspace. The lakehouse keeps copies of history, counts, reports, and file totals.
Two places save data:
- D1 is a database. It answers recent queries.
- Iceberg is files in the R2 bucket
workspace-events. It keeps the long history.
Some types go to both places. Then D1 is the hot store. Iceberg is the long copy. D1 keeps that copy for 28 days. A query about those 28 days reads D1. An older query reads Iceberg. Every 6 hours, the worker deletes D1 rows that are older than 28 days.
Cloudflare Pipelines writes the Iceberg files. The rows are ready after about 1 minute.
Document telemetry and Logic Pilot traces use both places. Document audit uses Iceberg only. The other types use D1 only.
Fields in every row
| Field | Meaning |
|---|---|
eventId | The unique id of this row. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the event happened. |
Most rows also carry companyId, the id of the company inside the workspace. Rows about one person also carry userId, and some also carry userEmail. Each data-type section below lists every field from the worker schema for that row.
See the route
This chart shows the same route, one type at a time. Choose a type. Drag the chart to move it. Use the controls to zoom.
Document audit
Document audit is the history of one document. It records who changed the document, and what changed.
The workspace sends the changes about 7 seconds after a burst of changes stops. A check every 5 minutes sends a change that the first listen missed.
The worker saves the row in Iceberg document_audit.v1 only.
The workspace reads this history when a person opens the audit trail or the activity of that document. Builder does not chart this type.
Example row (click to expand)
This is the full row from the worker schema.
{
"workspaceId": "d00232",
"eventId": "evt_audit_001",
"occurredAt": "2026-10-07T08:15:00.000Z",
"auditId": "audit_001",
"timestamp": "2026-10-07T08:15:00.000Z",
"documentTimestamp": "2026-10-07T08:14:58.000Z",
"serviceName": "logic-bee",
"type": "document",
"appId": "app_1",
"appSlug": "finance",
"appType": "workspace",
"projectId": "proj_1",
"deploymentId": "d00232",
"bobPackage": "finance",
"bobPackageVersion": "1.0.0",
"environment": "production",
"deployment": "d00232",
"licenseId": "lic_1",
"userId": "user_1",
"userEmail": "ada@example.com",
"roleType": "admin",
"userRoleId": "role_1",
"isSuperAdmin": "false",
"userType": "user",
"companyId": "co_1",
"businessUnitId": "bu_1",
"sessionId": "sess_1",
"ipAddress": "203.0.113.10",
"requestType": "http",
"userAgent": "Mozilla/5.0",
"complianceFlags": "[]",
"retentionPolicy": "default",
"retentionYears": 7,
"dataClassification": "internal",
"docId": "inv_1001",
"mongoId": "6640a1",
"docType": "invoice",
"namespace": "finance",
"docName": "invoice",
"cfpPath": "finance/invoice",
"status": "paid",
"version": 3,
"operationType": "update",
"operation": "update",
"jsonPatch": "[{\"op\":\"replace\",\"path\":\"/data/status\",\"value\":\"paid\"}]",
"patchOperationsCount": 1,
"fieldsAffected": "[\"/data/status\"]",
"updatedFields": "[\"/data/status\"]",
"removedFields": "[]",
"truncatedArrays": "[]",
"data": "{}",
"previousDocument": "{\"status\":\"unpaid\"}",
"currentDocument": "{\"status\":\"paid\"}",
"currentDocumentHash": "abc123",
"previousDocumentHash": "def456",
"message": "Status changed from unpaid to paid",
"activityMessage": "Ada marked the invoice paid",
"isActivity": true,
"userRequestId": "req_1",
"transactionId": "txn_1",
"podName": "logic-bee-1"
}
View all fields
| Field | Meaning |
|---|---|
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
eventId | The unique id of this row. |
occurredAt | The time the event happened. |
auditId | The id of the audit record. May be empty. |
timestamp | The time the audit record was written. |
documentTimestamp | The time stored on the document. |
serviceName | The service that wrote the change. |
type | The audit type. |
appId | The id of the app. |
appSlug | The slug of the app. |
appType | The type of the app. |
projectId | The id of the project. |
deploymentId | The id of the deployment. |
bobPackage | The package name. |
bobPackageVersion | The package version. |
environment | The environment, for example production. |
deployment | The deployment name. |
licenseId | The license id. |
userId | The id of the person who made the change. |
userEmail | The email of that person. |
roleType | The role type of that person. |
userRoleId | The id of that person's role. |
isSuperAdmin | Whether that person is a super admin, as a string. |
userType | The kind of user. |
companyId | The id of the company inside the workspace. |
businessUnitId | The id of the business unit. |
sessionId | The id of the session. |
ipAddress | The IP address of the request. |
requestType | The kind of request. |
userAgent | The user agent of the request. |
complianceFlags | A JSON string of compliance flags. |
retentionPolicy | The retention policy name. |
retentionYears | How many years the row is kept. |
dataClassification | The data classification label. |
docId | The id of the document that changed. |
mongoId | The database id of the document. |
docType | The type of the document. |
namespace | The namespace of the document. |
docName | The name of the document, for display. |
cfpPath | The CFP path of the document. |
status | The document status after the change. |
version | The document version after the change. |
operationType | The kind of operation. |
operation | The kind of change: create, update, or delete. |
jsonPatch | The exact change, as a JSON Patch array stored in a string. |
patchOperationsCount | The number of operations in jsonPatch. |
fieldsAffected | A JSON string of the paths the change touched. |
updatedFields | A JSON string of the paths that were updated. |
removedFields | A JSON string of the paths that were removed. |
truncatedArrays | A JSON string of arrays that were truncated. |
data | A JSON string of the change data. |
previousDocument | A JSON string of the document before the change. |
currentDocument | A JSON string of the document after the change. |
currentDocumentHash | A hash of the document after the change. |
previousDocumentHash | A hash of the document before the change. |
message | A short, human-readable summary of the change. |
activityMessage | The text shown in the activity list. |
isActivity | true when this row is also an activity item. |
userRequestId | The id of the user request. |
transactionId | The id of the transaction. |
podName | The name of the pod that wrote the change. |
The 5-minute check sends an audit change that the first listen missed. After the workspace prepares the history row, it does not send that change again.
Document telemetry
Document telemetry is a count of documents. It is not the documents.
Twice a day, each workspace sends the counts. Each row is one collection. count is the number of documents. countDeleted is the number of deleted documents.
The worker saves the same row in two places:
- D1
document_telemetry, tablev1, for 28 days - Iceberg
document_telemetry.v1
Builder reads these counts for the document charts. The overview of one deployment uses them too.
Example row (click to expand)
This is the full row from the worker schema.
{
"workspaceId": "d00232",
"eventId": "evt_tel_001",
"occurredAt": "2026-10-07T08:18:00.000Z",
"auditId": "",
"timestamp": "2026-10-07T08:18:00.000Z",
"serviceName": "logic-bee",
"type": "document-telemetry",
"appId": "app_1",
"appSlug": "finance",
"appType": "workspace",
"projectId": "proj_1",
"deploymentId": "d00232",
"bobPackage": "finance",
"bobPackageVersion": "1.0.0",
"environment": "production",
"deployment": "d00232",
"licenseId": "lic_1",
"userId": "user_1",
"userEmail": "ada@example.com",
"roleType": "admin",
"userRoleId": "role_1",
"isSuperAdmin": "false",
"userType": "user",
"companyId": "co_1",
"businessUnitId": "bu_1",
"sessionId": "sess_1",
"ipAddress": "203.0.113.10",
"requestType": "cron",
"userAgent": "",
"complianceFlags": "[]",
"retentionPolicy": "default",
"retentionYears": 7,
"dataClassification": "internal",
"collectionName": "finance",
"docType": "invoice",
"namespace": "finance",
"docName": "invoice",
"cfpPath": "finance/invoice",
"count": 1204,
"countDeleted": 16,
"version": 1
}
View all fields
| Field | Meaning |
|---|---|
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
eventId | The unique id of this row. |
occurredAt | The time the count was taken. |
auditId | May be empty. |
timestamp | The time the row was written. |
serviceName | The service that sent the count. |
type | The telemetry type. |
appId | The id of the app. |
appSlug | The slug of the app. |
appType | The type of the app. |
projectId | The id of the project. |
deploymentId | The id of the deployment. |
bobPackage | The package name. |
bobPackageVersion | The package version. |
environment | The environment, for example production. |
deployment | The deployment name. |
licenseId | The license id. |
userId | The id of the person, when the row has one. |
userEmail | The email of that person. |
roleType | The role type. |
userRoleId | The id of the role. |
isSuperAdmin | Whether the person is a super admin, as a string. |
userType | The kind of user. |
companyId | The id of the company inside the workspace. |
businessUnitId | The id of the business unit. |
sessionId | The id of the session. |
ipAddress | The IP address, when the row has one. |
requestType | The kind of request. |
userAgent | The user agent, when the row has one. |
complianceFlags | A JSON string of compliance flags. |
retentionPolicy | The retention policy name. |
retentionYears | How many years the row is kept. |
dataClassification | The data classification label. |
collectionName | The collection the count is for. |
docType | The type of document counted. |
namespace | The namespace of the collection. |
docName | The display name for the document type. |
cfpPath | The CFP path of the document type. |
count | The number of documents in the collection. |
countDeleted | The number of deleted documents in the collection. |
version | The version of this count row. |
The next twice-daily send sends a new snapshot.
Logic Pilot traces
A trace is one Logic Pilot run. A span is one step in that run.
The workspace waits until 512 spans are ready, or until 5 seconds pass. Then it sends the group. One group can contain spans from more than one run.
The worker saves the same row in two places:
- D1
mastra_spans, tablev2, for 28 days - Iceberg
mastra_spans.v2
Builder reads recent lists from D1. Builder reads one trace from Iceberg. A short cache holds the latest trace. A new save clears that cache.
Example row (click to expand)
This is the full row from the worker schema. endedAt is left out when the span has no end time. Do not send an empty string for it.
{
"workspaceId": "d00232",
"eventId": "evt_span_001",
"occurredAt": "2026-10-07T09:01:05.000Z",
"startAt": "2026-10-07T09:01:04.160Z",
"endedAt": "2026-10-07T09:01:05.000Z",
"traceId": "trace_88",
"spanId": "span_3",
"parentSpanId": "span_1",
"spanType": "model",
"name": "answer",
"model": "gpt-4.1",
"provider": "openai",
"durationMs": 840,
"promptTokens": 120,
"completionTokens": 40,
"status": "ok",
"companyId": "co_1",
"userId": "user_1",
"environment": "production",
"conversationId": "conv_9",
"messageId": "msg_9",
"mode": "text",
"inputJson": "{}",
"outputJson": "{}",
"attributesJson": "{}",
"metadataJson": "{}",
"errorJson": "",
"entityType": "conversation",
"entityId": "conv_9",
"entityName": "Invoice help",
"isEvent": false,
"isRootSpan": false,
"runId": "run_12",
"tags": "[]"
}
View all fields
| Field | Meaning |
|---|---|
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
eventId | The unique id of this row. |
occurredAt | The time the span was recorded. |
startAt | The time the span started. |
endedAt | The time the span ended. Omitted when there is no end time. |
traceId | The id of the run this span belongs to. |
spanId | The id of this step. |
parentSpanId | The id of the parent span. May be empty. |
spanType | The kind of step, for example model. |
name | A short label for the step. |
model | The AI model used in this step. May be empty. |
provider | The provider of that model, for example openai. May be empty. |
durationMs | How long the step took, in milliseconds. |
promptTokens | Input tokens used by this step. |
completionTokens | Output tokens used by this step. |
status | ok, or an error state. May be empty. |
companyId | The id of the company. May be empty. |
userId | The id of the person. May be empty. |
environment | The environment. May be empty. |
conversationId | The Logic Pilot conversation. May be empty. |
messageId | The message this span belongs to. May be empty. |
mode | The conversation mode. May be empty. |
inputJson | A JSON string of the span input. |
outputJson | A JSON string of the span output. |
attributesJson | A JSON string of span attributes. |
metadataJson | A JSON string of span metadata. |
errorJson | A JSON string of the error. Empty when there is no error. |
entityType | The kind of entity this span is about. May be empty. |
entityId | The id of that entity. May be empty. |
entityName | The name of that entity. May be empty. |
isEvent | true when this span is an event. |
isRootSpan | true for the first span in the run. |
runId | The id of the broader run that contains this trace. May be empty. |
tags | A JSON string of tags. |
A span group that fails to leave the workspace is dropped. Later spans go out in a new group.
Workspace stats
Workspace stats is a count of one company at one time. The count includes people, roles, companies, business units, invitations, API keys, two-factor use, and the autonomy level.
Twice a day, each workspace sends one row for each company. The worker saves it in D1 workspace_stats, table workspace_stats. If the same row is already saved, the worker leaves it.
Builder reads these rows to show change over time, and to show the latest counts.
Example row (click to expand)
This is the full row from the worker schema. eventId is workspaceId, companyId, and occurredAt.
{
"eventId": "d00232:co_1:2026-10-07T07:26:00.000Z",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T07:26:00.000Z",
"companyId": "co_1",
"users": 42,
"deletedUsers": 1,
"guests": 3,
"deletedGuests": 0,
"roles": 8,
"deletedRoles": 0,
"companies": 1,
"deletedCompanies": 0,
"businessUnits": 3,
"superAdminUsers": 1,
"adminUsers": 4,
"otherUsers": 37,
"apiKeys": 2,
"deletedApiKeys": 0,
"apiCredentials": 1,
"invitesJson": "{\"pending\":2}",
"guestInvitesJson": "{}",
"pendingInvites": 2,
"oldestPendingInviteAt": "2026-10-01T09:00:00.000Z",
"invitesTotal": 10,
"twoFactorRequired": 1,
"usersWith2fa": 30,
"rolesInUse": 6,
"autonomyLevel": 2
}
View all fields
| Field | Meaning |
|---|---|
eventId | workspaceId, companyId, and occurredAt. A repeat of the same row is left as it is. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the snapshot was taken. |
companyId | The id of the company. May be empty. |
users | The number of people in the company. |
deletedUsers | The number of deleted people. |
guests | The number of guests. |
deletedGuests | The number of deleted guests. |
roles | The number of roles defined. |
deletedRoles | The number of deleted roles. |
companies | The number of companies in the workspace. |
deletedCompanies | The number of deleted companies. |
businessUnits | The number of business units. |
superAdminUsers | The number of super admins. |
adminUsers | The number of admins. |
otherUsers | The number of people who are not admins. |
apiKeys | The number of API keys. |
deletedApiKeys | The number of deleted API keys. |
apiCredentials | The number of API credentials. |
invitesJson | A JSON string of invite counts by status. |
guestInvitesJson | A JSON string of guest invite counts by status. |
pendingInvites | The number of invitations not yet accepted. |
oldestPendingInviteAt | The time of the oldest pending invitation. May be empty. |
invitesTotal | The total number of invitations. |
twoFactorRequired | 1 if two-factor login is required, 0 if not. |
usersWith2fa | The number of people with two-factor login on. |
rolesInUse | The number of roles that are assigned. |
autonomyLevel | The autonomy level set for the company. |
The next twice-daily send sends a new row.
Logic Pilot stats
Logic Pilot stats is three counts. The workspace sends them twice a day.
- One row for each conversation that changed in the last 72 hours.
- One row of folder counts for each company.
- One row of token totals for each model.
The 72 hours are longer than the time between the two daily sends. If one send fails, the next send still includes that conversation.
The worker saves the rows in D1 logic_pilot_stats. The tables are logic_pilot_conversation_stat, logic_pilot_folder_stat, and token_telemetry_stat.
Builder reads them for the conversation charts and the token charts.
Conversation rows
One row for each conversation that changed in the last 72 hours.
Example row (click to expand)
This is the full row from the worker schema. eventId is workspaceId, conversationId, and updatedAt. hasContext, isBookmarked, and inFolder are stored as true or false.
{
"eventId": "d00232:conv_9:2026-10-06T18:00:00.000Z",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T10:43:00.000Z",
"updatedAt": "2026-10-06T18:00:00.000Z",
"createdAt": "2026-10-01T09:00:00.000Z",
"lastMessageAt": "2026-10-06T18:00:00.000Z",
"firstMessageAt": "2026-10-01T09:05:00.000Z",
"durationMs": 432000000,
"companyId": "co_1",
"userId": "user_1",
"roleType": "admin",
"conversationId": "conv_9",
"name": "Invoice help",
"nameLength": 12,
"status": "active",
"isBookmarked": false,
"inFolder": true,
"conversationType": "chat",
"mode": "text",
"modelId": "gpt-4.1",
"agentId": "agent_1",
"hasContext": true,
"contextItemCount": 2,
"eventsCount": 3,
"messageCount": 14,
"userMessageCount": 7,
"assistantMessageCount": 7,
"totalUserCharCount": 900,
"totalUserWordCount": 160,
"maxUserWordCount": 40,
"voiceMessageCount": 0,
"totalVoiceDurationMs": 0,
"avgVoiceDurationMs": 0,
"maxVoiceDurationMs": 0,
"documentCount": 1,
"spreadsheetCount": 0,
"chartCount": 0
}
View all fields
| Field | Meaning |
|---|---|
eventId | workspaceId, conversationId, and updatedAt. A repeat send is ignored. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the snapshot was taken. |
updatedAt | The time the conversation was last updated. May be empty. |
createdAt | The time the conversation was created. May be empty. |
lastMessageAt | The time of the last message. May be empty. |
firstMessageAt | The time of the first message. May be empty. |
durationMs | The length of the conversation, in milliseconds. |
companyId | The id of the company. May be empty. |
userId | The id of the person. May be empty. |
roleType | The role type of that person. May be empty. |
conversationId | The id of the conversation. |
name | The conversation name. May be empty. |
nameLength | The length of the name. |
status | active, or another conversation state. May be empty. |
isBookmarked | true when the conversation is bookmarked. |
inFolder | true when the conversation is in a folder. |
conversationType | The kind of conversation. May be empty. |
mode | text, or another conversation mode. May be empty. |
modelId | The AI model used in the conversation. May be empty. |
agentId | The id of the agent. May be empty. |
hasContext | true when the conversation has context items. |
contextItemCount | The number of context items. |
eventsCount | The number of events. |
messageCount | The number of messages in the conversation. |
userMessageCount | The number of messages from the person. |
assistantMessageCount | The number of messages from the assistant. |
totalUserCharCount | The total characters in the person's messages. |
totalUserWordCount | The total words in the person's messages. |
maxUserWordCount | The longest person message, in words. |
voiceMessageCount | The number of voice messages. |
totalVoiceDurationMs | The total voice length, in milliseconds. |
avgVoiceDurationMs | The average voice length, in milliseconds. |
maxVoiceDurationMs | The longest voice message, in milliseconds. |
documentCount | The number of documents produced. |
spreadsheetCount | The number of spreadsheets produced. |
chartCount | The number of charts produced. |
Folder rows
One row of folder counts for each company.
Example row (click to expand)
This is the full row from the worker schema. eventId is workspaceId, companyId, and occurredAt. One row is for one company, not for one folder.
{
"eventId": "d00232:co_1:2026-10-07T10:43:00.000Z",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T10:43:00.000Z",
"companyId": "co_1",
"folderCount": 6,
"emptyFolderCount": 1,
"usersWithFoldersCount": 4,
"maxFoldersPerUser": 3,
"avgFoldersPerUser": 1.5,
"conversationsInFoldersCount": 20,
"conversationsNotInFolderCount": 5,
"maxConversationsInFolder": 8,
"avgConversationsPerFolder": 3.3
}
View all fields
| Field | Meaning |
|---|---|
eventId | workspaceId, companyId, and occurredAt. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the snapshot was taken. |
companyId | The id of the company. May be empty. |
folderCount | The number of folders in the company. |
emptyFolderCount | The number of folders with no conversations. |
usersWithFoldersCount | The number of people who have at least one folder. |
maxFoldersPerUser | The most folders one person has. |
avgFoldersPerUser | The average number of folders per person. |
conversationsInFoldersCount | The number of conversations placed in a folder. |
conversationsNotInFolderCount | The number of conversations that are not in a folder. |
maxConversationsInFolder | The most conversations in one folder. |
avgConversationsPerFolder | The average number of conversations per folder. |
Token rows
One row of token totals for each model. These numbers are totals since the counter started, not just for that day.
Example row (click to expand)
This is the full row from the worker schema.
{
"eventId": "d00232:co_1:gpt-4.1:2026-10-07T10:43:00.000Z",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T10:43:00.000Z",
"companyId": "co_1",
"model": "gpt-4.1",
"input": 120000,
"output": 40000,
"cachedInput": 8000,
"total": 160000
}
View all fields
| Field | Meaning |
|---|---|
eventId | The unique id of this row. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the snapshot was taken. |
companyId | The id of the company. May be empty. |
model | The AI model these totals are for. |
input | Total input tokens sent to the model. |
output | Total output tokens returned by the model. |
cachedInput | Total input tokens served from cache. |
total | input plus output. |
The next twice-daily send sends a new snapshot. Conversation rows still cover the last 72 hours, so a missed send is not lost.
Bug reports
A person files a bug in the workspace. The workspace saves the bug first. Then it sends a copy to the worker.
The worker saves the copy in D1 feedbacks, table bug. Builder reads the copy.
Example row (click to expand)
This is the full row from the worker schema. A repeat uses bugDocId.
{
"eventId": "evt_bug_001",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T11:02:00.000Z",
"environment": "production",
"deploymentId": "d00232",
"licenseId": "lic_1",
"companyId": "co_1",
"companyName": "Northwind",
"userId": "user_1",
"userEmail": "ada@example.com",
"userName": "Ada",
"roleType": "admin",
"bugDocId": "bug_44",
"pageUrl": "https://workspace.example/invoices",
"screenshotUrl": "https://workspace.example/files/bug_44.png",
"reportType": "something-doesnt-work",
"description": "The invoice total does not update after a line change.",
"userTimezone": "Europe/Chisinau",
"date": "2026-10-07",
"userLanguage": "en"
}
View all fields
| Field | Meaning |
|---|---|
eventId | The unique id of this row. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the bug was filed. |
environment | The environment, for example production. |
deploymentId | The id of the deployment. |
licenseId | The license id. |
companyId | The id of the company. |
companyName | The name of the company. May be empty. |
userId | The id of the person who filed the bug. |
userEmail | The email of that person. |
userName | The name of that person. May be empty. |
roleType | The role type of that person. May be empty. |
bugDocId | The id of the bug document in the workspace. A repeat of the same bug uses this id. |
pageUrl | The page the person was on when they filed the bug. May be empty. |
screenshotUrl | The URL of the screenshot. May be empty. |
reportType | The kind of problem, for example something-doesnt-work. May be empty. |
description | The person's own description of the problem. May be empty. |
userTimezone | The person's time zone. May be empty. |
date | The date the person entered. May be empty. |
userLanguage | The person's language. May be empty. |
The bug report stays in the workspace. Only the Builder copy is missing.
Workspace feedback
A person sends feedback in the workspace. The workspace saves it first. Then it sends a copy to the worker.
The worker saves the copy in D1 feedbacks, table workspace_feedback. Builder reads the copy.
Example row (click to expand)
This is the full row from the worker schema. A repeat uses feedbackDocId.
{
"eventId": "evt_fb_001",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T11:10:00.000Z",
"environment": "production",
"deploymentId": "d00232",
"licenseId": "lic_1",
"companyId": "co_1",
"companyName": "Northwind",
"userId": "user_1",
"userEmail": "ada@example.com",
"userName": "Ada",
"roleType": "admin",
"feedbackDocId": "fb_12",
"workspaceUrl": "https://workspace.example",
"description": "The home page loads slowly in the morning.",
"date": "2026-10-07",
"userLanguage": "en",
"userTimezone": "Europe/Chisinau"
}
View all fields
| Field | Meaning |
|---|---|
eventId | The unique id of this row. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the feedback was sent. |
environment | The environment, for example production. |
deploymentId | The id of the deployment. |
licenseId | The license id. |
companyId | The id of the company. |
companyName | The name of the company. May be empty. |
userId | The id of the person who sent the feedback. |
userEmail | The email of that person. |
userName | The name of that person. May be empty. |
roleType | The role type of that person. May be empty. |
feedbackDocId | The id of the feedback document in the workspace. A repeat uses this id. |
workspaceUrl | The workspace the feedback came from. May be empty. |
description | The person's own description. May be empty. |
date | The date the person entered. May be empty. |
userLanguage | The person's language. May be empty. |
userTimezone | The person's time zone. May be empty. |
The feedback stays in the workspace. Only the Builder copy is missing.
Logic Pilot feedback
A person rates a Logic Pilot message. The rating stays on the session. When the session is saved, the workspace sends only new ratings.
The worker saves each new rating in D1 feedbacks, table logic_pilot_message. Builder reads the copy.
Example row (click to expand)
This is the full row from the worker schema. A repeat uses feedbackId. polarity is positive or negative.
{
"eventId": "evt_lp_001",
"workspaceId": "d00232",
"occurredAt": "2026-10-07T11:20:00.000Z",
"environment": "production",
"deploymentId": "d00232",
"licenseId": "lic_1",
"companyId": "co_1",
"companyName": "Northwind",
"userId": "user_1",
"userEmail": "ada@example.com",
"userName": "Ada",
"roleType": "admin",
"feedbackId": "rate_7",
"polarity": "negative",
"typeOfIssue": "wrong_data",
"details": "The answer used the wrong invoice.",
"sessionDocId": "sess_3",
"sessionName": "Invoice help",
"messageId": "msg_9",
"agentId": "agent_1",
"mode": "text",
"modelId": "gpt-4.1",
"messageContents": "[]",
"assistantVisualContents": "[]",
"conversationId": "conv_9"
}
View all fields
| Field | Meaning |
|---|---|
eventId | The unique id of this row. |
workspaceId | The id of the workspace that sent the row. This is the deployment id. |
occurredAt | The time the rating was saved. |
environment | The environment, for example production. |
deploymentId | The id of the deployment. |
licenseId | The license id. |
companyId | The id of the company. |
companyName | The name of the company. May be empty. |
userId | The id of the person who rated the message. |
userEmail | The email of that person. |
userName | The name of that person. May be empty. |
roleType | The role type of that person. May be empty. |
feedbackId | The id of this rating. A repeat uses this id. |
polarity | positive or negative. |
typeOfIssue | The kind of problem, for example wrong_data. May be empty. |
details | The person's own note about the rating. May be empty. |
sessionDocId | The id of the session document. |
sessionName | The name of the session. May be empty. |
messageId | The id of the rated message. |
agentId | The id of the agent. May be empty. |
mode | The conversation mode. May be empty. |
modelId | The AI model that produced the rated message. May be empty. |
messageContents | A JSON string of the message contents. |
assistantVisualContents | A JSON string of the assistant's visual contents. |
conversationId | The id of the conversation. May be empty. |
The rating stays on the session. Only the Builder copy is missing.
File storage
File storage is the size and the use of the workspace file buckets. The worker collects this data. The workspace does not send it.
Once a day, the worker asks Builder for deployments with status started or terminated. Then the worker reads the previous day's totals from Cloudflare. Each workspace has two buckets:
{workspaceId}-cdnis the public bucket{workspaceId}-storageis the private bucket
The worker saves one row for each bucket for that day in D1 workspace_stats, table workspace_r2_stat. A later save for the same day replaces that row. The worker reads the totals. It does not open the files.
Builder reads these rows for the storage charts.
Example row (click to expand)
This is the full row from the worker schema. eventId is bucketName, then |, then the UTC day. occurredAt is the start of that day. Storage fields are the highest value of the day. Operation and bandwidth fields are totals for the day.
{
"eventId": "d00232-cdn|2026-10-06",
"workspaceId": "d00232",
"occurredAt": "2026-10-06T00:00:00.000Z",
"bucketName": "d00232-cdn",
"bucketKind": "cdn",
"payloadSize": 1073741824,
"metadataSize": 1048576,
"objectCount": 540,
"uploadCount": 12,
"classAOps": 120,
"classBOps": 8000,
"freeOps": 10,
"errorOps": 2,
"clientErrorOps": 1,
"serverErrorOps": 1,
"opsByTypeJson": "{\"PutObject\":12}",
"bytesDownload": 524288000,
"bytesUpload": 10485760
}
View all fields
| Field | Meaning |
|---|---|
eventId | bucketName, then |, then the UTC day YYYY-MM-DD. |
workspaceId | The id of the workspace. This is the deployment id. |
occurredAt | The start of that UTC day. |
bucketName | The bucket this row is for. |
bucketKind | cdn for the public bucket, storage for the private bucket. |
payloadSize | The highest total bytes stored in the bucket that day. |
metadataSize | The highest metadata bytes stored that day. |
objectCount | The highest number of files in the bucket that day. |
uploadCount | The number of uploads that day. |
classAOps | Count of write-type operations that day. |
classBOps | Count of read-type operations that day. |
freeOps | Count of free operations that day. |
errorOps | Count of error operations that day. |
clientErrorOps | Count of client error operations that day. |
serverErrorOps | Count of server error operations that day. |
opsByTypeJson | A JSON string of request counts by Cloudflare action type. |
bytesDownload | Total bytes downloaded that day. |
bytesUpload | Total bytes uploaded that day. |
This type carries no companyId or userId.
That day's group is skipped. A later save for the same day replaces the row.
Quick map
| Data | Sent | Where it is saved | Kept for | Who reads it |
|---|---|---|---|---|
| Document audit | On each change, about 7 seconds after a burst stops | Iceberg document_audit.v1 | Kept in Iceberg; no deletion rule given | The workspace |
| Document telemetry | Twice a day | D1 for 28 days, and Iceberg document_telemetry.v1 | 28 days in D1; kept in Iceberg after that | Builder |
| Logic Pilot traces | In groups, as spans are ready | D1 for 28 days, and Iceberg mastra_spans.v2 | 28 days in D1; kept in Iceberg after that | Builder |
| Workspace stats | Twice a day | D1 workspace_stats | Kept in D1; no deletion rule given | Builder |
| Logic Pilot stats | Twice a day | D1 logic_pilot_stats | Kept in D1; no deletion rule given | Builder |
| Bug reports | When a person files a bug | D1 feedbacks, table bug | Kept in D1; no deletion rule given | Builder |
| Workspace feedback | When a person sends feedback | D1 feedbacks, table workspace_feedback | Kept in D1; no deletion rule given | Builder |
| Logic Pilot feedback | When a person rates a message | D1 feedbacks, table logic_pilot_message | Kept in D1; no deletion rule given | Builder |
| File storage | Once a day (the worker reads it) | D1 workspace_r2_stat | Kept in D1; no deletion rule given | Builder |
Data that is not in the lakehouse
Builder activity, build plans, and CFPs stay in Builder. Their charts read Builder.
The documents, sessions, and files that a person works on stay in the workspace.
Resources
This architecture needs three programs and the Cloudflare resources below.
The workspace backend already exists. It is logic-bee. It runs in each workspace and sends the data.
bob-api already exists. It is the Builder backend. It reads the lakehouse for charts.
The worker is bob-workspace-orchestration. It receives the data and saves it. Create that worker with the Cloudflare resources in the table below.
Programs
| Program | Repository | What it does |
|---|---|---|
| Workspace backend | logic-bee | Sends data from each workspace. This program already exists. |
| bob-api | bob-api | Reads the lakehouse for Builder charts. This program already exists. |
| Worker | bob-workspace-orchestration | Receives the data and saves it. Its Cloudflare resources are in the next table. |
Cloudflare
This table is the Cloudflare resources. Use these names. Do not invent new ones.
Iceberg types use a namespace and a version. The Iceberg table is {namespace}.{version}. The stream is {namespace}_{version}. The sink is {namespace}_{version}_sink. The pipeline is {namespace}_{version}_pipeline.
When that type also has a hot store, the D1 database name is the namespace and the D1 table name is the version (v1 or v2).
A D1-only type uses a database name and a table name that describe the data. Several tables can share one database.
| Name | Provider | Type | Binding | Parent | Data | Documentation |
|---|---|---|---|---|---|---|
bob-workspace-orchestration | Cloudflare | Worker | — | — | All lakehouse types | Workers |
workspace-events | Cloudflare | R2 bucket | WORKSPACE_EVENTS | — | Iceberg files | R2 |
workspace_events_cache | Cloudflare | KV namespace | WORKSPACE_EVENTS_CACHE | bob-workspace-orchestration | Document audit and Logic Pilot traces | KV |
document_telemetry | Cloudflare | D1 database | DOCUMENT_TELEMETRY | bob-workspace-orchestration | Document telemetry | D1 |
mastra_spans | Cloudflare | D1 database | MASTRA_SPANS | bob-workspace-orchestration | Logic Pilot traces | D1 |
feedbacks | Cloudflare | D1 database | FEEDBACKS | bob-workspace-orchestration | Bug reports, workspace feedback, and Logic Pilot feedback | D1 |
logic_pilot_stats | Cloudflare | D1 database | LOGIC_PILOT_STATS | bob-workspace-orchestration | Logic Pilot stats | D1 |
workspace_stats | Cloudflare | D1 database | WORKSPACE_STATS | bob-workspace-orchestration | Workspace stats and file storage | D1 |
v1 | Cloudflare | D1 table | — | document_telemetry | Document telemetry | D1 |
v2 | Cloudflare | D1 table | — | mastra_spans | Logic Pilot traces | D1 |
bug | Cloudflare | D1 table | — | feedbacks | Bug reports | D1 |
workspace_feedback | Cloudflare | D1 table | — | feedbacks | Workspace feedback | D1 |
logic_pilot_message | Cloudflare | D1 table | — | feedbacks | Logic Pilot feedback | D1 |
logic_pilot_conversation_stat | Cloudflare | D1 table | — | logic_pilot_stats | Logic Pilot stats | D1 |
logic_pilot_folder_stat | Cloudflare | D1 table | — | logic_pilot_stats | Logic Pilot stats | D1 |
token_telemetry_stat | Cloudflare | D1 table | — | logic_pilot_stats | Logic Pilot stats | D1 |
workspace_stats | Cloudflare | D1 table | — | workspace_stats | Workspace stats | D1 |
workspace_r2_stat | Cloudflare | D1 table | — | workspace_stats | File storage | D1 |
document_audit.v1 | Cloudflare | Iceberg table | — | workspace-events | Document audit | R2 Data Catalog |
document_telemetry.v1 | Cloudflare | Iceberg table | — | workspace-events | Document telemetry | R2 Data Catalog |
mastra_spans.v2 | Cloudflare | Iceberg table | — | workspace-events | Logic Pilot traces | R2 Data Catalog |
document_audit_v1 | Cloudflare | Pipeline stream | DOCUMENT_AUDIT_STREAM | bob-workspace-orchestration | Document audit | Pipelines |
document_audit_v1_sink | Cloudflare | Pipeline sink | — | workspace-events | Document audit | Pipelines |
document_audit_v1_pipeline | Cloudflare | Pipeline | — | — | Document audit | Pipelines |
document_telemetry_v1 | Cloudflare | Pipeline stream | DOCUMENT_TELEMETRY_STREAM | bob-workspace-orchestration | Document telemetry | Pipelines |
document_telemetry_v1_sink | Cloudflare | Pipeline sink | — | workspace-events | Document telemetry | Pipelines |
document_telemetry_v1_pipeline | Cloudflare | Pipeline | — | — | Document telemetry | Pipelines |
mastra_spans_v2 | Cloudflare | Pipeline stream | MASTRA_SPANS_STREAM | bob-workspace-orchestration | Logic Pilot traces | Pipelines |
mastra_spans_v2_sink | Cloudflare | Pipeline sink | — | workspace-events | Logic Pilot traces | Pipelines |
mastra_spans_v2_pipeline | Cloudflare | Pipeline | — | — | Logic Pilot traces | Pipelines |
0 2 * * * | Cloudflare | Cron | — | bob-workspace-orchestration | File storage | Cron Triggers |
0 */6 * * * | Cloudflare | Cron | — | bob-workspace-orchestration | Document telemetry and Logic Pilot traces | Cron Triggers |
Each workspace has its own file buckets. The worker only reads their totals. Those buckets are not in this list.
The worker vars DOCUMENT_AUDIT_ICEBERG_TABLE, DOCUMENT_TELEMETRY_ICEBERG_TABLE, and MASTRA_SPANS_ICEBERG_TABLE must match the three Iceberg table names.
Cloudflare limits
These limits are set by Cloudflare. They apply to the worker, the stores, the pipelines, and the reads on this page. The numbers were taken from the Cloudflare docs on 7 October 2026. Where a doc lists a Free plan and a Paid plan, the table below uses the Paid plan. The Free plan is lower. Pipelines and R2 SQL are in open beta, so those limits can change.
Worker
The worker receives every ingest and runs both crons. Workers limits.
| Limit | Paid plan | What it means here |
|---|---|---|
| Memory | 128 MB per isolate | A large ingest batch or a large query result must not be held in memory all at once. |
| CPU time, HTTP | 30 seconds by default, up to 5 minutes | Waiting on D1, KV, R2, or fetch does not count as CPU time. |
| CPU time, cron shorter than 1 hour | 30 seconds | This page has no cron that runs that often. |
| CPU time, cron of 1 hour or longer | 15 minutes | The file-storage cron and the 28-day cleanup use this limit. |
| Cron wall time | 15 minutes | A cron that is still running after 15 minutes is stopped. |
| Subrequests | 10,000 per invocation | Each call to D1, KV, R2, or fetch counts. |
| Connections waiting for headers | 6 at one time | A seventh call waits until one of the six receives headers. |
| Request body | 100 MB on Free and Pro, 200 MB on Business, up to 5 GB on Enterprise | A body over the zone limit is rejected with HTTP 413. |
waitUntil | 30 seconds after the response | Work left running after the response can be stopped after 30 seconds. |
An account can have 250 cron triggers on the Paid plan, and 5 on the Free plan. A cron change can take up to 15 minutes to reach the whole network. Cron time is UTC. Cron Triggers.
D1
D1 is the hot store. Each database is separate. D1 limits.
| Limit | Paid plan | What it means here |
|---|---|---|
| Database size | 10 GB | Cloudflare does not raise this. Five databases are used: document_telemetry, mastra_spans, feedbacks, logic_pilot_stats, and workspace_stats. |
| Queries per invocation | 1,000 | One cron run shares this cap across the databases it touches. |
| SQL duration | 30 seconds | A delete of a huge set of rows in one statement can hit this. The 28-day cleanup deletes in batches. |
| Row size | 2 MB | One row, including JSON strings, must stay under 2 MB. |
| SQL statement | 100 KB | |
| Bound parameters | 100 per query | |
| Columns per table | 100 | |
| Concurrency | One query at a time per database | Extra queries wait. A full queue returns an overloaded error. |
Pipelines
Pipelines carry document audit, document telemetry, and Logic Pilot traces into Iceberg. The product is in open beta. Pipelines limits.
| Limit | Value | What it means here |
|---|---|---|
| Streams, sinks, and pipelines | 20 of each per account | This page uses 3 of each. |
| Ingest request | 5 MB | One send to a stream must stay under 5 MB. |
| Ingest rate | 5 MB per second per stream | |
| Stream schema | Fixed after creation | A new column needs a new stream, sink, pipeline, and Iceberg table. Logic Pilot traces did this for v2. |
| Bad events | Dropped while processing | An event that does not match the schema is accepted, then dropped. |
| Iceberg sink | Parquet only | A sink cannot be attached to an Iceberg table that already exists. |
| Roll interval | 60 seconds minimum | Rows are ready about 1 minute after the worker sends them. |
Streams. R2 Data Catalog sink.
R2 and Iceberg
Iceberg files live in the R2 bucket workspace-events. R2 limits. Table maintenance.
| Limit | Value | What it means here |
|---|---|---|
| Bucket storage and object count | No stated maximum | |
| One object | 5 TiB | |
| Writes to the same object name | 1 per second | |
| Catalog jurisdiction | Default jurisdiction only | A bucket in the EU or FedRAMP jurisdiction cannot use the catalog. |
| Compaction file size | 64 MB to 512 MB | Compaction reads Parquet files only. Files that no snapshot points at are not deleted. |
| Snapshot expiration | Off unless enabled | If it is enabled, the default is to drop snapshots older than 30 days and keep the last 5. |
R2 SQL
Builder and the workspace read Iceberg through R2 SQL. R2 SQL is read-only and in open beta. It reads Parquet only. R2 SQL limitations. R2 SQL troubleshooting.
| Limit | Value | What it means here |
|---|---|---|
| Writes | Not allowed | New rows arrive through Pipelines, not through R2 SQL. |
LIMIT | 10,000 rows maximum | |
OFFSET | Not supported | A later page uses WHERE and ORDER BY, not OFFSET. |
| JSON inside a string column | No JSON path filter | jsonPatch and the other JSON strings are filtered in the application, or stored as their own columns. |
| Heavy queries | Can time out or return HTTP 400 | A large COUNT(DISTINCT), a large sort, or a join of three big tables can be rejected. |
KV
workspace_events_cache holds the short cache for a document audit and for one Logic Pilot trace. KV limits.
| Limit | Paid plan | What it means here |
|---|---|---|
| Writes to the same key | 1 per second | A burst of saves for one trace can hit this. |
| Operations per invocation | 1,000 | |
| Value size | 25 MiB | |
| Key size | 512 bytes |
GraphQL Analytics
The file-storage cron reads bucket totals from the Cloudflare GraphQL Analytics API. It does not read the files. GraphQL API limits.
| Limit | Value | What it means here |
|---|---|---|
| Requests | 300 per 5 minutes for one user or token | The cron runs once a day, so this cap is wide for that job. |
| Accounts in one query | 1 |