I originally published this article on the Microsoft Tech Community.
Overview of GCP log types and SOC value
Modern Security Operations Centers (SOCs) require visibility into key Google Cloud Platform logs to detect threats and suspicious activities. The main log types include:
- GCP Audit Logs — These encompass Admin Activity, Data Access, and Access Transparency logs for GCP services. They record every administrative action (resource creation, modification, IAM changes, etc.) and access to sensitive data, providing a trail of who did what and when in the cloud. In a SOC context, audit logs help identify unauthorized changes or anomalous admin behavior (e.g. an attacker creating a new service account or disabling logging). They are essential for compliance and forensics, detailing changes to configurations and access patterns across GCP resources.
- VPC Flow Logs — These capture network traffic flow information at the Virtual Private Cloud (VPC) level. Each entry typically includes source/destination IPs, ports, protocol, bytes, and an allow/deny action. In a SOC, VPC flow logs are invaluable for network threat detection: monitoring access patterns, detecting port scanning, identifying unusual internal traffic, and profiling ingress/egress traffic for anomalies. A surge in outbound traffic to an unknown IP or lateral movement between VMs can be spotted via flow logs. They also aid in investigating data exfiltration and verifying network policy enforcement.
- Cloud DNS Logs — Google Cloud DNS query logs record DNS requests/responses from resources, and DNS audit logs record changes to DNS configurations. Query logs are extremely useful in threat hunting because they can reveal systems resolving malicious domains (C2 servers, phishing sites, DGA domains) or performing unusual lookups. Many malware campaigns rely on DNS, so having these logs in Sentinel enables detection of known-bad domains and anomalous DNS patterns. DNS audit logs track modifications to DNS records (e.g. newly added subdomains or changed IP mappings), which can indicate misconfigurations or domain-hijacking attempts.
Together, these GCP logs provide comprehensive coverage: audit logs tell you what actions were taken in the cloud, while VPC and DNS logs tell you what network activities are happening. Ingesting all three into Sentinel gives a cloud security architect visibility to detect unauthorized access, network intrusions, and malware communication in a GCP environment.
Native Microsoft Sentinel GCP connector: architecture & setup
Microsoft Sentinel offers native data connectors to ingest Google Cloud logs, leveraging Google’s Pub/Sub messaging for scalable, secure integration. The native solution is built on Sentinel’s Codeless Connector Framework (CCF) and uses a pull-based architecture: GCP exports logs to Pub/Sub, and Sentinel’s connector pulls from Pub/Sub into Azure. This approach avoids custom code and uses cloud-native services on both sides.
Supported GCP log connectors. Out of the box, Sentinel provides connectors for:
- GCP Audit Logs — the Cloud Audit Logs (admin activity, data access, transparency).
- GCP Security Command Center (SCC) — security findings from Google SCC for threat and vulnerability management.
- GCP VPC Flow Logs — (recently added) VPC network flow logs.
- GCP Cloud DNS Logs — (recently added) Cloud DNS query logs and DNS audit logs.
- Others — additional connectors for specific GCP services (Cloud Load Balancer logs, Cloud CDN, Cloud IDS, GKE, IAM activity, etc.) via CCF. Each connector typically writes to its own Log Analytics table (e.g.
GCPAuditLogs,GCPVPCFlow,GCPDNS) and comes with built-in KQL parsers.
Architecture & authentication. The native connector uses Google Pub/Sub as the pipeline for log delivery. On the Google side you set up a Pub/Sub topic that receives the logs (via Cloud Logging exports), and Sentinel subscribes to that topic. Authentication is handled through Workload Identity Federation (WIF) using OpenID Connect: instead of managing static credentials, you establish trust between Azure AD and GCP so Sentinel can impersonate a GCP service account. The high-level flow is:
GCP Cloud Logging (logs from services) → Log Router (export sink) → Pub/Sub Topic → (secure pull over OIDC) → Azure Sentinel Data Connector → Log Analytics Workspace.
This ensures a secure, keyless integration. The Azure side authenticates as a Google service account via OIDC tokens issued by Azure AD, which GCP trusts through the Workload Identity Provider.
GCP setup (publishing logs to Pub/Sub)
-
Enable required APIs. Ensure the GCP project hosting the logs has the IAM API and Cloud Resource Manager API enabled (needed for creating identity pools and roles). You’ll also need owner/editor access on the project to create the resources below.
-
Create a Workload Identity Pool & Provider. In Google Cloud IAM, create a new Workload Identity Pool (e.g.
Azure-Sentinel-Pool) and then a Workload Identity Provider within it that trusts your Azure AD tenant. Google provides a template for Azure AD OIDC trust — you supply your Azure tenant ID and the audience and issuer URIs Azure uses. For Azure Commercial, the issuer is typicallyhttps://sts.windows.net/<TenantID>/and audienceapi://<some-guid>as documented. -
Create a service account for Sentinel. Still in GCP, create a dedicated service account. This is what Sentinel “impersonates” via the WIF trust. Grant it two key roles:
- Pub/Sub Subscriber on the subscription that will be created (allows pulling messages). You can grant
roles/pubsub.subscriberat the project level or on the specific subscription. - Workload Identity User on the pool. Add a principal of the form
principalSet://iam.googleapis.com/projects/<WIF_project_number>/locations/global/workloadIdentityPools/<Pool_ID>/*and grant itroles/iam.workloadIdentityUseron your service account, allowing the Azure AD identity to impersonate it.
Note: GCP best practice is often to keep the identity pool in a centralized project and service accounts in separate projects, but the Sentinel connector UI has expected them in one project. It’s simplest to create the WIF pool/provider and the service account within the same GCP project unless docs confirm cross-project support.
- Pub/Sub Subscriber on the subscription that will be created (allows pulling messages). You can grant
-
Create a Pub/Sub topic and subscription. Create a Topic (e.g.
projects/yourproject/topics/sentinel-logs) — one topic per log type is convenient. Add a Subscription in Pull mode (e.g.sentinel-audit-sub). Defaults are fine, or extend retention if you want messages to persist longer during downtime (default is 7 days). -
Create a logging export (sink). In Cloud Logging → Logs Router, create a sink:
- Give it a descriptive name (e.g.
audit-logs-to-sentinel). - Choose Cloud Pub/Sub as the destination and select your topic.
- Scope and filters: decide which logs to include. For audit logs you might include all audit logs in the project; for VPC/DNS use an inclusion filter for those specific log names (e.g.
logName:"compute.googleapis.com/vpc_flows"). Organization-level sinks can aggregate logs from all projects. - Permissions: when creating the sink, GCP asks to grant the sink service account publish rights to the topic. Accept so logs can flow.
Verify logs are flowing: in Pub/Sub → Subscriptions, “Pull” messages manually. Generating a test event (create a VM to produce an audit log, or make a DNS query) helps confirm.
- Give it a descriptive name (e.g.
At this point GCP is set up to export logs. Google also provides Terraform scripts (and Microsoft supplies Terraform templates on GitHub) to automate the IAM and Pub/Sub configuration quickly.
Azure Sentinel setup (connecting the GCP connector)
- Install the GCP solution. In your Sentinel workspace, under Content Hub (or Data Connectors), find e.g. Google Cloud Platform Audit Logs and click Install. Repeat for VPC Flow / DNS as needed.
- Open data connector configuration. Go to Data Connectors, search for “GCP Pub/Sub Audit Logs,” select it, click Open connector page, then + Add new to configure a connection instance.
- Enter GCP parameters. Supply the Project ID, Project Number, the Topic and Subscription name, and the Service Account Email you created. Also enter the Workload Identity Provider ID (format
projects/<proj>/locations/global/workloadIdentityPools/<pool>/providers/<provider>). A common error is mixing up project ID (name) with project number, or using the wrong Tenant ID. - Data Collection Rule (DCR). Newer CCF connectors use the Log Ingestion API, so a DCR is used behind the scenes. If prompted, provide a name (docs often suggest prefixing with
Microsoft-Sentinel-, e.g.Microsoft-Sentinel-GCPAuditLogs-DCR). The system creates the DCR and a DCE for you. - Connect. Sentinel verifies the subscription exists and that the service account can authenticate. Errors like “WIF Pool ID not found” or “subscription not found” indicate a mismatch in IDs or permissions.
- Validation. After ~15–30 minutes, run a Log Analytics query such as
GCPAuditLogs | take 5(orGCPVPCFlow | take 5). Enable the data-connector health feature for alerts on latency and volume.
Data flow and ingestion
Once connected, the system works continuously and serverlessly:
- GCP Log Router pushes new log entries to Pub/Sub as they occur.
- The Sentinel connector polls the Pub/Sub subscription, using the service account credentials (via OIDC token) to pull messages in batches, typically every few seconds.
- Each message (a JSON log entry) is ingested into Log Analytics. The CCF connector uses the Log Ingestion API, mapping JSON fields to the table schema.
- Sentinel’s built-in parsers or Normalized Schemas (ASIM) let you query the logs in a friendly way.
This native pipeline is fully managed — no servers or code to run. Pub/Sub and OIDC make it scalable and secure by design.
Design considerations & best practices for the native connector
- Scalability & performance. Pub/Sub handles very large log volumes with low latency, and the CCF connectors use a SaaS, auto-scaling model. In testing, the native connector reliably ingests millions of log events per day.
- Reliability. Pub/Sub’s at-least-once delivery makes the integration robust — messages buffer during transient outages and the connector catches up on the backlog. It acknowledges messages only after ingestion, preventing data loss. Use connector health metrics to catch issues.
- Security. Workload Identity Federation eliminates persistent credentials — no service-account key file to leak. Give the service account minimal roles (essentially Pub/Sub subscription access). The Azure AD app underpinning the connector only needs Sentinel workspace access.
- Filtering and log-volume management. Filter GCP logs at the sink to avoid ingesting superfluous data (e.g. exclude noisy Data Access reads). Consider multiple sinks to separate log types, since each Sentinel connector ties to one subscription and one table.
- Coverage gaps. Check what the native connectors currently support. If a needed log type isn’t supported (e.g. a custom application writing to Cloud Logging), consider the custom approach below.
- Monitoring and troubleshooting. Each configured connector instance shows a status and last-received timestamp. On GCP, monitor the Pub/Sub subscription backlog — a healthy system has near-zero unacked messages.
- Multi-project or org-wide ingestion. Deploy a connector per project, or use an organization-level sink to funnel logs from all projects into a single Pub/Sub. The Terraform script supports an org sink via an organization ID.
In summary, the native GCP connector provides a straightforward, robust way to get Google Cloud logs into Sentinel — the recommended approach for supported log types due to minimal maintenance and tight integration.
Custom ingestion architecture (Pub/Sub to Azure Event Hub, etc.)
When the built-in connector doesn’t meet requirements — unsupported log types, custom formats, or a policy to use an intermediary — you can design a custom pipeline. One reference pattern is GCP Pub/Sub → Azure Event Hub → Sentinel.
- GCP export (source). Same as the native setup: create log sinks that export to Pub/Sub topics, with subscriptions for your pipeline to pull from. A push subscription can call an HTTP endpoint directly, but a pull model is more common for custom solutions.
- Bridge / transfer component — the core of the pipeline, a piece of code that reads from Pub/Sub and sends data to Azure. Options:
- Google Cloud Function / Cloud Run (in GCP): trigger on new Pub/Sub messages, parse, then forward to Azure. Keeps the pull logic on the GCP side and scales out automatically.
- Custom puller in Azure (Function App or container): use Google’s Pub/Sub client library to pull messages (with a service-account key), then send to Log Analytics. Centralizes everything in Azure but requires managing GCP credentials there.
- Google Cloud Dataflow (Apache Beam): a heavy-duty streaming option for very large scale with exactly-once processing — usually overkill unless you already use Beam.
- Destination in Azure — two primary ways to ingest into Sentinel:
- Azure Log Ingestion API (via DCR): the modern method. Create a DCE and a DCR that route incoming data to a Log Analytics table (e.g.
GCP_Custom_Logs_CL). Your bridge calls the ingestion REST endpoint with an Azure AD token, and the DCR can transform fields. This replaces the older HTTP Data Collector API. - Azure Event Hub + Sentinel connector: use an Event Hub as an intermediate buffer. Either use Sentinel’s Event Hub connector (expects a known format like CEF/JSON), or an Azure Function with an Event Hub trigger that calls the Log Ingestion API. Useful if you want a cloud-neutral queue between GCP and Azure, or to fan out to other systems.
- Azure Log Ingestion API (via DCR): the modern method. Create a DCE and a DCR that route incoming data to a Log Analytics table (e.g.
- Data transformation. With a custom pipeline you own any transformations:
- Message decoding: Pub/Sub messages hold the log entry base64-encoded in the
datafield; decode it to get the raw JSON. - Schema mapping: map important JSON fields to table columns (e.g.
src_ip,dest_ip,bytes_sent,action) via the DCR transformation. - Enrichment (optional): e.g. IP geolocation or threat-intel tagging — keep it light to avoid failure points.
- Filtering: drop noisy or benign events at the bridge to save cost.
- Batching: the Log Ingestion API accepts many records per call (up to ~1 MB / 30,000 events); batch for throughput.
- Message decoding: Pub/Sub messages hold the log entry base64-encoded in the
Authentication & permissions (custom pipeline). You handle two hops:
- GCP → bridge: a Cloud Function triggered by Pub/Sub gets the message directly; pulling from Azure needs a service-account key with minimal permissions (Pub/Sub subscriber only), stored securely (Key Vault or encrypted app setting) and rotated periodically.
- Bridge → Azure: create an Azure AD app registration and grant it Monitoring Metrics/Data roles (e.g. Monitoring Data Contributor) or the DCR’s ingestion role, then use the client ID/secret to get a token. Alternatively use a DCR SAS token. Follow least privilege — ingest only, no read of other data.
End-to-end example (firewall/VPC logs). A GCP log sink filters VPC firewall logs to a Pub/Sub topic. An Azure Function on a timer pulls messages each minute (authenticating with a stored service-account key), decodes each JSON payload, maps fields to a custom table (GCPFirewall_CL), and POSTs a batch to the Log Ingestion API using an Azure AD app’s credentials. It acknowledges Pub/Sub only after a successful send, so failures are retried — near-real-time with only seconds of latency.
Tooling. Use official SDKs (Google Pub/Sub client libraries; Azure Monitor Ingestion SDK). Use Terraform/IaC for repeatability. Implement robust logging (Cloud Logging on GCP, Application Insights on Azure) and alerts on failures or dropped volume.
In short, the custom route requires more effort up front but grants complete control — tune what you collect, transform to your needs, and integrate logs not natively supported. A hybrid approach (native for audit logs, custom for a niche source) is also common.
Comparison of native vs. custom ingestion
| Aspect | Native GCP connector (Pub/Sub integration) | Custom ingestion pipeline (Event Hub / API) |
|---|---|---|
| Ease of setup | Low-code: configure GCP and Azure, no custom code. Terraform scripts + a UI wizard; enable in a few hours if prerequisites are met. | High-code: design and write integration code plus configure services. Longer setup — days to weeks to develop and test. |
| Log-type coverage | Standard GCP logs (audit, SCC, VPC flow, DNS, etc.), limited to released connectors. | Virtually any log that can reach Pub/Sub, including custom application logs — you build parsing per format. |
| Development & maintenance | Minimal: runs as a managed service; Microsoft handles updates. | Ongoing: you own the code, credentials, and monitoring — closer to software maintenance. |
| Scalability | Cloud-scale and managed; auto-scales behind the scenes. | Depends on your implementation; you configure concurrency, throughput units, batching. |
| Reliability & resilience | Reliable by design (Pub/Sub durability + managed ingestion, built-in retries and health). | Varies; you implement retry/error handling and health checks. |
| Flexibility & transformation | Standardized ingestion into predefined tables; little transformation. | Highly flexible — choose fields, formats, enrichment, filtering, table layout. |
| Cost | No connector charge; minimal Pub/Sub cost plus Log Analytics ingestion. No infrastructure to run. | Adds Function/Event Hub execution and egress costs, but filtering can reduce Log Analytics spend. |
| Support & troubleshooting | Vendor-supported; connector UI surfaces errors; Microsoft ships updates. | Self-supported; debugging spans two clouds. |
Use the native connector whenever possible — it’s easier and reliably maintained. Opt for custom only when you truly need the flexibility or must support logs the native connectors can’t handle. Some teams start custom out of necessity and later migrate to native to cut maintenance.
Troubleshooting common issues
No data appearing in Sentinel. Be patient (~10–30 minutes for initial data). If still nothing:
- Check the Log Router sink status in GCP — correct inclusion filters? Are logs actually being generated (view them in Cloud Logging)?
- In Pub/Sub, use “Pull” on the subscription. If you can pull manually but Sentinel isn’t receiving, the issue is on the Azure side.
- Ensure the connector shows Connected. A single typo in Project Number or Service Account will block ingestion.
- Auth/connectivity errors (native): “Workload Identity Pool ID not found” or “Subscription does not exist” mean mismatched values. Double-check the Workload Identity Provider ID (including project number), the service account email and its Workload Identity User role, and the subscription name plus
roles/pubsub.subscriber. Ensure the Microsoft multi-tenant Sentinel connector app isn’t blocked in your tenant.
Tip: if you set things up manually and it’s failing, running Microsoft’s Terraform scripts can pinpoint what’s missing.
Partial data or specific logs missing. Revisit the sink filter (too narrow?). For DNS you may need both _Default and DNS-specific log IDs; for audit logs, Data Access logs must be enabled per service in GCP. For SCC, enable continuous export of findings. Check whether data landed in a differently-named table.
Duplicate logs (custom ingestion). If your function crashes after sending to Sentinel but before acking Pub/Sub, messages get retried and double-counted. Acknowledge only after successful ingestion, or deduplicate using a unique ID (audit logs have insertId).
Log Ingestion API errors (custom pipeline).
- 400 Bad Request — schema mismatch; the JSON doesn’t match the DCR (wrong type, missing/extra column).
- 403 Forbidden — auth failure; refresh the token or fix role assignments.
- 429 Too Many Requests — throttling; back off, retry, and batch more.
- Function timeouts — increase the timeout or split work into smaller chunks.
Connector health alerts. If you enabled health alerts (“no logs received in X minutes”), confirm whether it’s genuine (no new events in GCP) or a real pipeline break, then investigate accordingly.
Updating or migrating pipelines. When replacing a custom pipeline with a native connector (or vice versa), avoid running both at once or you’ll double-ingest. Plan a cutover, note the data may land in a different table, and adjust queries/workbooks — backfilling history if continuity matters.
By following these practices and monitoring closely, you get a centralized view in Sentinel where Azure, AWS, on-prem, and GCP logs all reside — empowering your SOC to run advanced detections and investigations across a multi-cloud environment from a single pane of glass.