← All posts

Ingesting Google Cloud Logs into Microsoft Sentinel: Native vs. Custom Architectures

Bringing GCP audit, VPC flow, and DNS logs into Microsoft Sentinel — the native Pub/Sub connector versus a custom ingestion pipeline, with setup steps, trade-offs, and troubleshooting.

I originally published this article on the Microsoft Tech Community.

Overview of GCP log types and SOC value

Modern Security Operations Centers (SOCs) require visibility into key Google Cloud Platform logs to detect threats and suspicious activities. The main log types include:

Together, these GCP logs provide comprehensive coverage: audit logs tell you what actions were taken in the cloud, while VPC and DNS logs tell you what network activities are happening. Ingesting all three into Sentinel gives a cloud security architect visibility to detect unauthorized access, network intrusions, and malware communication in a GCP environment.

Native Microsoft Sentinel GCP connector: architecture & setup

Microsoft Sentinel offers native data connectors to ingest Google Cloud logs, leveraging Google’s Pub/Sub messaging for scalable, secure integration. The native solution is built on Sentinel’s Codeless Connector Framework (CCF) and uses a pull-based architecture: GCP exports logs to Pub/Sub, and Sentinel’s connector pulls from Pub/Sub into Azure. This approach avoids custom code and uses cloud-native services on both sides.

Supported GCP log connectors. Out of the box, Sentinel provides connectors for:

Architecture & authentication. The native connector uses Google Pub/Sub as the pipeline for log delivery. On the Google side you set up a Pub/Sub topic that receives the logs (via Cloud Logging exports), and Sentinel subscribes to that topic. Authentication is handled through Workload Identity Federation (WIF) using OpenID Connect: instead of managing static credentials, you establish trust between Azure AD and GCP so Sentinel can impersonate a GCP service account. The high-level flow is:

GCP Cloud Logging (logs from services) → Log Router (export sink) → Pub/Sub Topic → (secure pull over OIDC) → Azure Sentinel Data Connector → Log Analytics Workspace.

This ensures a secure, keyless integration. The Azure side authenticates as a Google service account via OIDC tokens issued by Azure AD, which GCP trusts through the Workload Identity Provider.

GCP setup (publishing logs to Pub/Sub)

  1. Enable required APIs. Ensure the GCP project hosting the logs has the IAM API and Cloud Resource Manager API enabled (needed for creating identity pools and roles). You’ll also need owner/editor access on the project to create the resources below.

  2. Create a Workload Identity Pool & Provider. In Google Cloud IAM, create a new Workload Identity Pool (e.g. Azure-Sentinel-Pool) and then a Workload Identity Provider within it that trusts your Azure AD tenant. Google provides a template for Azure AD OIDC trust — you supply your Azure tenant ID and the audience and issuer URIs Azure uses. For Azure Commercial, the issuer is typically https://sts.windows.net/<TenantID>/ and audience api://<some-guid> as documented.

  3. Create a service account for Sentinel. Still in GCP, create a dedicated service account. This is what Sentinel “impersonates” via the WIF trust. Grant it two key roles:

    • Pub/Sub Subscriber on the subscription that will be created (allows pulling messages). You can grant roles/pubsub.subscriber at the project level or on the specific subscription.
    • Workload Identity User on the pool. Add a principal of the form principalSet://iam.googleapis.com/projects/<WIF_project_number>/locations/global/workloadIdentityPools/<Pool_ID>/* and grant it roles/iam.workloadIdentityUser on your service account, allowing the Azure AD identity to impersonate it.

    Note: GCP best practice is often to keep the identity pool in a centralized project and service accounts in separate projects, but the Sentinel connector UI has expected them in one project. It’s simplest to create the WIF pool/provider and the service account within the same GCP project unless docs confirm cross-project support.

  4. Create a Pub/Sub topic and subscription. Create a Topic (e.g. projects/yourproject/topics/sentinel-logs) — one topic per log type is convenient. Add a Subscription in Pull mode (e.g. sentinel-audit-sub). Defaults are fine, or extend retention if you want messages to persist longer during downtime (default is 7 days).

  5. Create a logging export (sink). In Cloud Logging → Logs Router, create a sink:

    • Give it a descriptive name (e.g. audit-logs-to-sentinel).
    • Choose Cloud Pub/Sub as the destination and select your topic.
    • Scope and filters: decide which logs to include. For audit logs you might include all audit logs in the project; for VPC/DNS use an inclusion filter for those specific log names (e.g. logName:"compute.googleapis.com/vpc_flows"). Organization-level sinks can aggregate logs from all projects.
    • Permissions: when creating the sink, GCP asks to grant the sink service account publish rights to the topic. Accept so logs can flow.

    Verify logs are flowing: in Pub/Sub → Subscriptions, “Pull” messages manually. Generating a test event (create a VM to produce an audit log, or make a DNS query) helps confirm.

At this point GCP is set up to export logs. Google also provides Terraform scripts (and Microsoft supplies Terraform templates on GitHub) to automate the IAM and Pub/Sub configuration quickly.

Azure Sentinel setup (connecting the GCP connector)

  1. Install the GCP solution. In your Sentinel workspace, under Content Hub (or Data Connectors), find e.g. Google Cloud Platform Audit Logs and click Install. Repeat for VPC Flow / DNS as needed.
  2. Open data connector configuration. Go to Data Connectors, search for “GCP Pub/Sub Audit Logs,” select it, click Open connector page, then + Add new to configure a connection instance.
  3. Enter GCP parameters. Supply the Project ID, Project Number, the Topic and Subscription name, and the Service Account Email you created. Also enter the Workload Identity Provider ID (format projects/<proj>/locations/global/workloadIdentityPools/<pool>/providers/<provider>). A common error is mixing up project ID (name) with project number, or using the wrong Tenant ID.
  4. Data Collection Rule (DCR). Newer CCF connectors use the Log Ingestion API, so a DCR is used behind the scenes. If prompted, provide a name (docs often suggest prefixing with Microsoft-Sentinel-, e.g. Microsoft-Sentinel-GCPAuditLogs-DCR). The system creates the DCR and a DCE for you.
  5. Connect. Sentinel verifies the subscription exists and that the service account can authenticate. Errors like “WIF Pool ID not found” or “subscription not found” indicate a mismatch in IDs or permissions.
  6. Validation. After ~15–30 minutes, run a Log Analytics query such as GCPAuditLogs | take 5 (or GCPVPCFlow | take 5). Enable the data-connector health feature for alerts on latency and volume.

Data flow and ingestion

Once connected, the system works continuously and serverlessly:

This native pipeline is fully managed — no servers or code to run. Pub/Sub and OIDC make it scalable and secure by design.

Design considerations & best practices for the native connector

  1. Scalability & performance. Pub/Sub handles very large log volumes with low latency, and the CCF connectors use a SaaS, auto-scaling model. In testing, the native connector reliably ingests millions of log events per day.
  2. Reliability. Pub/Sub’s at-least-once delivery makes the integration robust — messages buffer during transient outages and the connector catches up on the backlog. It acknowledges messages only after ingestion, preventing data loss. Use connector health metrics to catch issues.
  3. Security. Workload Identity Federation eliminates persistent credentials — no service-account key file to leak. Give the service account minimal roles (essentially Pub/Sub subscription access). The Azure AD app underpinning the connector only needs Sentinel workspace access.
  4. Filtering and log-volume management. Filter GCP logs at the sink to avoid ingesting superfluous data (e.g. exclude noisy Data Access reads). Consider multiple sinks to separate log types, since each Sentinel connector ties to one subscription and one table.
  5. Coverage gaps. Check what the native connectors currently support. If a needed log type isn’t supported (e.g. a custom application writing to Cloud Logging), consider the custom approach below.
  6. Monitoring and troubleshooting. Each configured connector instance shows a status and last-received timestamp. On GCP, monitor the Pub/Sub subscription backlog — a healthy system has near-zero unacked messages.
  7. Multi-project or org-wide ingestion. Deploy a connector per project, or use an organization-level sink to funnel logs from all projects into a single Pub/Sub. The Terraform script supports an org sink via an organization ID.

In summary, the native GCP connector provides a straightforward, robust way to get Google Cloud logs into Sentinel — the recommended approach for supported log types due to minimal maintenance and tight integration.

Custom ingestion architecture (Pub/Sub to Azure Event Hub, etc.)

When the built-in connector doesn’t meet requirements — unsupported log types, custom formats, or a policy to use an intermediary — you can design a custom pipeline. One reference pattern is GCP Pub/Sub → Azure Event Hub → Sentinel.

  1. GCP export (source). Same as the native setup: create log sinks that export to Pub/Sub topics, with subscriptions for your pipeline to pull from. A push subscription can call an HTTP endpoint directly, but a pull model is more common for custom solutions.
  2. Bridge / transfer component — the core of the pipeline, a piece of code that reads from Pub/Sub and sends data to Azure. Options:
    • Google Cloud Function / Cloud Run (in GCP): trigger on new Pub/Sub messages, parse, then forward to Azure. Keeps the pull logic on the GCP side and scales out automatically.
    • Custom puller in Azure (Function App or container): use Google’s Pub/Sub client library to pull messages (with a service-account key), then send to Log Analytics. Centralizes everything in Azure but requires managing GCP credentials there.
    • Google Cloud Dataflow (Apache Beam): a heavy-duty streaming option for very large scale with exactly-once processing — usually overkill unless you already use Beam.
  3. Destination in Azure — two primary ways to ingest into Sentinel:
    • Azure Log Ingestion API (via DCR): the modern method. Create a DCE and a DCR that route incoming data to a Log Analytics table (e.g. GCP_Custom_Logs_CL). Your bridge calls the ingestion REST endpoint with an Azure AD token, and the DCR can transform fields. This replaces the older HTTP Data Collector API.
    • Azure Event Hub + Sentinel connector: use an Event Hub as an intermediate buffer. Either use Sentinel’s Event Hub connector (expects a known format like CEF/JSON), or an Azure Function with an Event Hub trigger that calls the Log Ingestion API. Useful if you want a cloud-neutral queue between GCP and Azure, or to fan out to other systems.
  4. Data transformation. With a custom pipeline you own any transformations:
    • Message decoding: Pub/Sub messages hold the log entry base64-encoded in the data field; decode it to get the raw JSON.
    • Schema mapping: map important JSON fields to table columns (e.g. src_ip, dest_ip, bytes_sent, action) via the DCR transformation.
    • Enrichment (optional): e.g. IP geolocation or threat-intel tagging — keep it light to avoid failure points.
    • Filtering: drop noisy or benign events at the bridge to save cost.
    • Batching: the Log Ingestion API accepts many records per call (up to ~1 MB / 30,000 events); batch for throughput.

Authentication & permissions (custom pipeline). You handle two hops:

End-to-end example (firewall/VPC logs). A GCP log sink filters VPC firewall logs to a Pub/Sub topic. An Azure Function on a timer pulls messages each minute (authenticating with a stored service-account key), decodes each JSON payload, maps fields to a custom table (GCPFirewall_CL), and POSTs a batch to the Log Ingestion API using an Azure AD app’s credentials. It acknowledges Pub/Sub only after a successful send, so failures are retried — near-real-time with only seconds of latency.

Tooling. Use official SDKs (Google Pub/Sub client libraries; Azure Monitor Ingestion SDK). Use Terraform/IaC for repeatability. Implement robust logging (Cloud Logging on GCP, Application Insights on Azure) and alerts on failures or dropped volume.

In short, the custom route requires more effort up front but grants complete control — tune what you collect, transform to your needs, and integrate logs not natively supported. A hybrid approach (native for audit logs, custom for a niche source) is also common.

Comparison of native vs. custom ingestion

Aspect Native GCP connector (Pub/Sub integration) Custom ingestion pipeline (Event Hub / API)
Ease of setup Low-code: configure GCP and Azure, no custom code. Terraform scripts + a UI wizard; enable in a few hours if prerequisites are met. High-code: design and write integration code plus configure services. Longer setup — days to weeks to develop and test.
Log-type coverage Standard GCP logs (audit, SCC, VPC flow, DNS, etc.), limited to released connectors. Virtually any log that can reach Pub/Sub, including custom application logs — you build parsing per format.
Development & maintenance Minimal: runs as a managed service; Microsoft handles updates. Ongoing: you own the code, credentials, and monitoring — closer to software maintenance.
Scalability Cloud-scale and managed; auto-scales behind the scenes. Depends on your implementation; you configure concurrency, throughput units, batching.
Reliability & resilience Reliable by design (Pub/Sub durability + managed ingestion, built-in retries and health). Varies; you implement retry/error handling and health checks.
Flexibility & transformation Standardized ingestion into predefined tables; little transformation. Highly flexible — choose fields, formats, enrichment, filtering, table layout.
Cost No connector charge; minimal Pub/Sub cost plus Log Analytics ingestion. No infrastructure to run. Adds Function/Event Hub execution and egress costs, but filtering can reduce Log Analytics spend.
Support & troubleshooting Vendor-supported; connector UI surfaces errors; Microsoft ships updates. Self-supported; debugging spans two clouds.

Use the native connector whenever possible — it’s easier and reliably maintained. Opt for custom only when you truly need the flexibility or must support logs the native connectors can’t handle. Some teams start custom out of necessity and later migrate to native to cut maintenance.

Troubleshooting common issues

No data appearing in Sentinel. Be patient (~10–30 minutes for initial data). If still nothing:

Tip: if you set things up manually and it’s failing, running Microsoft’s Terraform scripts can pinpoint what’s missing.

Partial data or specific logs missing. Revisit the sink filter (too narrow?). For DNS you may need both _Default and DNS-specific log IDs; for audit logs, Data Access logs must be enabled per service in GCP. For SCC, enable continuous export of findings. Check whether data landed in a differently-named table.

Duplicate logs (custom ingestion). If your function crashes after sending to Sentinel but before acking Pub/Sub, messages get retried and double-counted. Acknowledge only after successful ingestion, or deduplicate using a unique ID (audit logs have insertId).

Log Ingestion API errors (custom pipeline).

Connector health alerts. If you enabled health alerts (“no logs received in X minutes”), confirm whether it’s genuine (no new events in GCP) or a real pipeline break, then investigate accordingly.

Updating or migrating pipelines. When replacing a custom pipeline with a native connector (or vice versa), avoid running both at once or you’ll double-ingest. Plan a cutover, note the data may land in a different table, and adjust queries/workbooks — backfilling history if continuity matters.

By following these practices and monitoring closely, you get a centralized view in Sentinel where Azure, AWS, on-prem, and GCP logs all reside — empowering your SOC to run advanced detections and investigations across a multi-cloud environment from a single pane of glass.