Splunk Cloud Platform · migration planning On-prem collection tier sizing Paths per Splunk documentation

Cloud ingest architecture planner

Splunk Cloud stores and searches your data, but it cannot collect all of it. Some sources reach Splunk Cloud with nothing for you to build. Others need a collector server on your own network. Answer a few questions below and this planner tells you which is which, how many servers you would need to build, what to ask your firewall team to open, and what your team needs to do.

Planning estimate only. Component counts are starting points for an architecture review, not a validated design. Throughput per collector varies with event size, parsing complexity, and network path. Validate against a pilot. See the full disclaimer.

How to use this planner

  1. Answer the highlighted questions. Fields with a green bar and an Answer tag are the ones that change your result. Example numbers are pre-filled, so replace them with yours. A best guess is fine, and you can change it any time.
  2. Leave anything tagged Default OK alone. These are expert tuning settings, tucked into collapsed “Advanced settings” boxes. Ignore them for a first estimate.
  3. Tick boxes tagged If it applies only for things you actually use.
  4. Read the results on the right, top to bottom. They update as you type. Download the plan to share with your team or with Splunk.
Numbers worth having ready (about 10 minutes)
  • How much data you send to Splunk per day (GB/day). Your Splunk license page or your Splunk admin will know.
  • How many servers and PCs will run the Splunk agent (a “universal forwarder”), and roughly what share of your data they produce.
  • If you collect network device logs: about how many syslog events per second they send. Your network team can tell you.
  • Which cloud and SaaS services you want in Splunk (Microsoft 365, AWS, Okta and so on), and roughly how much data they produce.
  • Whether any data sources cannot reach the internet, and whether your Splunk Cloud stack is FedRAMP.

Unfamiliar word? See the plain-English glossary at the bottom of the page.

1. Your Splunk Cloud stack

The basics: how much data you send, and how many places you collect it from.

How much data you send (or expect to send) to Splunk Cloud each day. Find it on your Splunk license page, or ask your Splunk admin. Not sure? A best guess is fine.

The number of separate data centers or office networks where you would place collection servers. Most customers: 1.

Splunk decides which one your stack uses. Not sure? Leave Victoria. To check: in Splunk Cloud open Support & Services > About. It only changes whether Splunk runs an extra server (the “IDM”) for cloud and SaaS sources.

Advanced settings defaults are fine for a first estimate

Peak factor: how much busier your busiest period is than an average moment. 2.5 means peaks carry 2.5 times the average rate. Compression: how small data gets on the trip to Splunk Cloud. 50 means it shrinks by half. It only affects the bandwidth estimate.

2. Servers and PCs (forwarders)

A universal forwarder is a small Splunk agent installed on a server or PC. It reads logs there and sends them to Splunk.

How many servers and PCs will send logs, and roughly what percent of your daily data they produce. A rough guess is fine. If you are not using agents at all, enter 0 and 0.

Leave on “Straight to Splunk Cloud” unless your firewall rules force traffic through a few controlled servers. Splunk prefers the direct route.

Advanced settings only used if agents go through your servers

Planning assumptions for the intermediate servers (heavy forwarders). These are field-practice figures, not Splunk limits. Replace them only if you have measured your own.

3. Network devices (syslog and flow)

Syslog is how routers, firewalls, switches and appliances send logs. Flow data (NetFlow) records who talked to whom. Splunk Cloud cannot receive either directly, so something on your network must collect them.

Ask your network team or check your current syslog server. Rough guide: a small network is a few hundred, a mid-size one 1,500 to 5,000, and a large busy one can be 10,000 or more.

Not sure? Pick SC4S. It is free from Splunk and already knows how to label many device types. Pick “our own syslog servers” only if you already run rsyslog or syslog-ng and want to keep them.

Each kind of device needs its own labelling rules when you run your own servers. Files are kept on disk in case the connection to Splunk drops.

Advanced settings defaults are fine for a first estimate

Message size turns events per second into GB per day. 350 bytes is typical. One-server capacity: 9,000 is the most an untuned SC4S handled over UDP with no loss in Splunk’s test. Splunk says to test your own hardware.

4. Cloud services and applications

Cloud and SaaS services (Microsoft 365, AWS, Okta, ServiceNow and similar) are polled by Splunk Cloud itself, so they need no server of yours. Applications and scripts can also push data to Splunk over the internet, using an endpoint called HEC.

How many cloud and SaaS services you want in Splunk, and roughly what percent of your daily data they produce. A rough guess is fine. These never touch your internet link or your servers.

Advanced settings defaults are fine for a first estimate

5. Managing agents and reducing data

Optional extras. Most customers with more than a few dozen agents need a way to manage them, and many want to send less data.

Advanced settings defaults are fine for a first estimate

10,000 is a cautious planning figure. Splunk documents up to 25,000 per deployment server, in a cluster of up to three (75,000 agents in total), but does not publish a separate limit for one standalone server. Busy fleets that check in often can reach their practical limit sooner. Not sure? Leave 10,000. Count every Splunk instance that checks in, including heavy forwarders.

Your result at a glance Victoria

Servers to build is the number of machines your team must install, patch and monitor. Anything Splunk hosts for you is not counted. Sizes are planning estimates, so confirm them with a small pilot.

Servers to build
—
—
Data needing a collector
—
—
Data going straight to Splunk
—
no collector of yours needed
Data indexed after filtering
—
—
Peak internet bandwidth
—
compressed, at peak

Where your daily data comes from

Your total, split by source. This should add up to your total. If it does not, adjust the shares on the left.

SourcePer dayShare

Ingest paths

One row per kind of source: where it starts, what carries it, and where it lands. Amber boxes are things you build on your network. Green boxes need nothing of yours.

Servers you would build

The machines behind the “Servers to build” number. Cores and RAM are per server. Sizes follow Splunk guidance where it exists and are starting points elsewhere.

ComponentQtyCoresRAMDiskWhy it exists

Splunk-hosted components not your servers

Things Splunk runs for you. Shown so you know they exist, but you do not build or patch them.

ComponentRun byWhat it does for this plan

How common data sources get to Splunk typical sources

Find your own sources in the list. Amber in the last column means something on your network must collect it. Green means Splunk Cloud collects it or it goes straight in.

Kind of dataHow it is collectedWho collects it

Work items —

The tasks your team will need to do, in order. The tag shows who usually owns each one: sysops (server admins), netops (network team), infosec (security team) or splunk (Splunk admins). Technical detail is included for those teams. Use “copy all” to paste the list into a ticket.

Firewall rules to request

The connections your firewalls need to allow. “From” is the side that starts the connection. Hand this table to your network team.

FromToPortPurpose

Things to know

Green is good news, amber means watch out, red means fix before relying on the numbers, and blue is good to know.

Plain-English glossary
Universal forwarder
A small Splunk agent installed on a server or PC. It reads logs on that machine and sends them to Splunk. Also called a UF.
Heavy forwarder
A full Splunk server that receives data, can process it, and passes it on. Used as a middleman or to collect things agents cannot. Also called an HF.
Syslog
A very common way for network devices and appliances to send logs. Usually over UDP, which has no delivery guarantee.
Events per second (EPS)
How many log messages arrive each second. Used to size syslog collectors.
SC4S
Splunk Connect for Syslog. A free Splunk-supported syslog collector that labels many device types automatically and sends data to Splunk Cloud over HEC.
HEC
HTTP Event Collector. A web address on Splunk Cloud where applications and scripts can send data with a token, with no agent needed.
Splunk Stream
The Splunk add-on that collects network flow data (NetFlow, IPFIX, sFlow) on your network and forwards it to Splunk.
Victoria and Classic
The two “Experiences” of Splunk Cloud. Splunk chooses yours and expects everyone to move to Victoria. Victoria runs cloud and SaaS collection for you on its own servers.
IDM
Inputs Data Manager. On Classic only, a Splunk-run server that polls cloud and SaaS sources for you. It cannot take HEC data or syslog.
Add-on
A package from Splunkbase that knows how to collect and label one kind of data, such as Microsoft 365 or Okta.
Deployment server
A server that pushes settings and add-ons to all your agents from one place. Splunk Cloud does not run one for you. For large fleets or high availability, Splunk supports a cluster of up to three deployment servers behind a load balancer.
Edge Processor
A Splunk tool that runs on your network and filters or masks data before it is sent, so it never uses your internet link or counts toward your data allowance. Configured from Splunk Cloud.
Ingest Actions
A Splunk Cloud feature that filters or masks data after it arrives but before it is indexed. No servers of yours needed.
High availability (HA) and N+1
Keeping collection running if a server fails. Here it means one extra server per tier, so the tier keeps its full capacity after losing one.
Peak factor
How much busier your busiest period is than an average moment. 2.5 means peaks carry 2.5 times the average rate.
WAN egress and bandwidth
Data leaving your network toward Splunk Cloud. Data that Splunk Cloud collects itself from cloud services never uses your internet link.
Collector
Any server of yours that gathers data and sends it on to Splunk Cloud, such as SC4S, a heavy forwarder or a Stream server.
Credentials package
A file you download from your Splunk Cloud stack that lets your agents connect securely to it.
ACS
Admin Config Service. A Splunk Cloud interface for automating admin tasks, such as private connectivity or IP allow lists.
CIM
Common Information Model. A standard set of field names that lets dashboards and detections work across many data sources.

Sources for the rules applied above

  1. Victoria runs modular and scripted inputs on search heads with no IDM, and Classic needs one. Splunk assigns the Experience, and Victoria also allows self-service install of private and most Splunkbase apps. Determine your Splunk Cloud Platform Experience.
  2. The IDM does not support HEC inputs, is not a syslog sink, and cannot receive unencrypted TCP. Introduction to Getting Data In. Forwarder ingest uses 9997 and HEC uses 443: Splunk Cloud Platform Experiences.
  3. Intermediate forwarders: size from peak volume and sender count, and avoid funnelling many inputs into too few. The two-pipelines-per-indexer guideline is waived behind a load balancer, which Splunk says is the case on Victoria. Intermediate data routing.
  4. SC4S: the SC4S team gives no general throughput estimates. In its test on 16 vCPU, default settings sustained about 14,600 msg/s on one TCP connection and about 81,000 across ten, and UDP was clean at 4,500 and 9,000 EPS but lost about half the messages at 27,000 until tuned. SC4S performance tests. Its architecture guidance is to scale vertically, avoid load balancing syslog, and collect in the same VLAN as the devices: SC4S architecture. Disk buffer sizing follows SC4S configuration.
  5. Edge Processor needs Splunk Cloud 9.0.2209 or later in a supported region, runs on customer-managed Linux hosts, listens on 9997 for forwarders and 8088 for HEC by default (syslog has no default port), and is sized by daily volume. About the Edge Processor solution, installation requirements, sizing guidelines.
  6. Ingest Actions filters, masks and routes data in Splunk Cloud from a search head. Use ingest actions.
  7. DB Connect uses JDBC. On Victoria it can run on a search head once outbound ports are opened through the ACS API. Deploy DB Connect to Victoria.
  8. Splunk Stream forwarders send to HEC, support NetFlow v5 and v9, IPFIX, sFlow and jFlow, and have no default flow port. Ingest NetFlow and IPFIX with Stream.
  9. Deployment server: Splunk documents up to 25,000 clients per deployment server in a cluster of up to three (75,000 in total). This planner defaults to a more cautious 10,000 per server, which you can raise up to 25,000. A cluster needs Splunk Enterprise 9.2 or later on each server, a shared drive and a load balancer or DNS name: Implement a deployment server cluster. Above 50 clients it must be a dedicated Splunk Enterprise instance, the reference host is 12 cores and 12 GB, and above 10,000 clients Splunk advises tuning dedicatedIoThreads: Estimate deployment server performance. Large forwarder fleets can also use a tool such as Puppet or Chef, and heavy forwarders that do more than the forwarder license covers need a license from Support: Use forwarders to get data into Splunk Cloud Platform.
  10. Certificate auto-rotation works for forwarders 9.3.0 or later against Splunk Cloud 9.2.2406 or later, and only when they connect directly to the stack. Enable a receiver for Splunk Cloud Platform.
  11. Private connectivity covers forwarder, HEC and optionally search traffic. About private connectivity.
  12. The Cloud Monitoring Console has forwarder dashboards but does not cover your own SC4S, heavy forwarder or Edge Processor hosts. Use the Forwarder dashboards.
  13. Heavy forwarder throughput, syslog capacity per node, and wire compression are planning assumptions, not published Splunk figures. Each is an input so you can substitute measured numbers.

Planning a cloud migration? August Schell has migrated federal Splunk environments to Splunk Cloud end to end, including collection tier design, add-on remediation, and forwarder cutover.

Talk to an architect