Cloud & Automation
Report Engine Architecture: Request to Result
Trace a report request through gateway, security, validation, cache, Fabric query, result construction and monitoring.
On this page
Report Engine Architecture: Request to Result is a production engineering concern, not just an implementation detail. The useful design connects user intent to controlled execution, measurable outcomes, and evidence that remains available during an incident.
Request lifecycle
API Management
Step 1: Receive request
Assign a request ID, accept a correlation ID
Report Engine
Step 2: Validate request
Report, format, date range, row and byte limits
Report Engine
Step 3: Resolve tenant & security
Tenant from trusted claims; authorize report and filters
Storage account
Step 4: Check cache
Tenant-aware key; fresh output → skip to step 6 or 8
Microsoft Fabric
Step 5: Query Fabric if needed
Approved template, bound parameters, timeout
Only on a cache miss
Report Engine
Step 6: Build result
Shape rows into JSON, CSV or Parquet
Storage account
Step 7: Store output / cache result
Write to the tenant’s folder with a TTL
Report Engine
Step 8: Return response
JSON body or short-lived download link
The client reaches API Management or a gateway, which establishes identity, request size limits, and coarse throttling. The report engine performs report authorization, tenant validation, parameter validation, and finer workload limits before any query work begins.
Cache decision
- RequestValidated, tenant resolved
- Cache key
tenantreportparametersformatversion - Look upIn
report-cache, under{tenant}/{report}/… - Cached and fresh?Within the report’s TTL
Yes Cache hit
- Authorize again for this caller
- Read the stored output
- Return result
No Cache miss
- Query Fabric (approved template)
- Build output: JSON, CSV or Parquet
- Store under the tenant’s path, with a TTL
- Return result
Tenant-aware: same report, same parameters, different tenants
- contoso
contoso · sales-by-region · 2026-09 · csv · v3report-cache/contoso/sales-by-region/… - fabrikam
fabrikam · sales-by-region · 2026-09 · csv · v3report-cache/fabrikam/sales-by-region/…
Two keys, two folders: never shared.
Build a canonical cache key from trusted tenant identity, report type, normalized date range, dimensions, filters, format, and report version. The same safe logical request produces the same key. A hit still requires authorization. A miss proceeds to the controlled query layer and may cache the finished result with TTL and freshness metadata.
Query and result
The query layer selects an approved query template for Fabric Warehouse or Lakehouse SQL endpoint and binds validated values. The result builder enforces row and byte limits, produces JSON, CSV, Parquet or another allowed format, and stores large artifacts privately. Download URLs are short lived and checked against tenant identity.
Observability
Record request_id, correlation_id, tenant_id, client_id, api_key_id, report_name, requested_at, started_at, completed_at, duration_ms, cache_hit, query_duration_ms, rows_returned, output_bytes, status, and error_code. These fields connect directly to the Monitoring pillar and support capacity, reliability, audit, and cost analysis.
Production considerations
Treat tenant isolation as a first-class boundary. Derive tenant identity from authenticated claims or a trusted service mapping, never only from tenant_id supplied by a browser. Apply tenant filters in the controlled query layer, include tenant identity in cache keys, protect stored results, and write an audit event for access. Redact secrets and sensitive filter values from ordinary logs.
Every operation needs a request or execution ID plus a correlation ID that crosses service boundaries. Capture UTC timestamps, status, duration, workload size, and error code. Keep high-cardinality detail in logs or traces rather than unbounded metric labels. Make telemetry asynchronous and bounded so a monitoring outage cannot take down the production path.
Define limits before scale exposes missing policy: maximum date range, maximum rows and bytes, execution timeout, concurrency per tenant, queue capacity, retry budget, and artifact retention. Reject invalid work early with a specific response. Retry only transient operations and use idempotency keys where duplicate execution could create extra files or charges.
Validate with representative data and failure drills. Test empty results, boundary dates, invalid dimensions, cross-tenant attempts, dependency timeouts, cache corruption, cancellation, retries, and large outputs. Compare the visible result with source totals and retain enough context to reproduce the decision. Operational readiness means an on-call engineer can identify the failing layer and take a bounded action without guessing.
Tradeoffs and failure modes
More telemetry improves diagnosis but adds storage, privacy, and cardinality costs. More caching reduces query load but creates freshness and invalidation risks. More flexible requests improve usefulness but expand the security and performance surface. Prefer explicit report or execution contracts, allow-listed variation, and measured exceptions over an unrestricted interface.
Watch for partial success: a pipeline can write data and fail before logging completion; a report can finish after its caller disconnects; a cache write can fail after a valid result was returned. Model those outcomes explicitly. Do not relabel an unknown value as zero, and do not overwrite failed attempts when a retry succeeds. Preserve the original error even if secondary logging also fails.
Practical rollout
Begin with one important workload and a small set of service objectives. Instrument the complete path, establish a baseline, and review evidence with application, data, security, and operations owners. Add alerts only when the receiver has a documented response. Expand by workload class after identifiers, access controls, and retention have proved reliable. This creates an operating model, not merely a dashboard.
Operational review questions
Before release, ask whether a responder can identify the affected client, tenant, workload class, code version, and dependency from retained evidence. Confirm that success means the intended data was delivered, not merely that a process exited without an exception. Check whether a retry is safe, whether cancellation stops downstream work, and whether partial output can be mistaken for a complete result.
During review, compare normal, peak, and failure behavior with representative volume. Verify that limits produce explicit outcomes and that dashboards distinguish rejected, queued, executing, completed, failed, and cancelled work. Assign ownership for the service, data contract, alerts, cache or operational store, and recovery procedure. Record decisions close to the implementation so future changes preserve the reasoning. Finally, test the investigation path with someone who did not build the feature; if that person cannot move from symptom to a specific execution and dependency, the design still lacks operational context.