Skip to content
Eudora Technology
Home  / Insights
Cloud Technology

Is Serverless Right for Your Application?

Indian Railway 139 Server Room.
Photo: 139 Server Room 01 by Indrajit Das, CC BY-SA 3.0, via Wikimedia Commons

Serverless is the right call when your traffic is uneven and your code finishes quickly. If your application sits idle most of the day and then gets busy for two hours, you will pay less and patch nothing. If it runs flat out around the clock, or a single request takes several minutes, you will pay more than a plain virtual machine would cost and you will fight the platform’s limits while doing it. Everything below is the detail behind that one-paragraph answer.

CSIRO's latest supercomputer cluster will be among the world's first to combine traditional CPUs with more powerful graphics processing units or GPUs, providing a world class compu
Managed runtimes remove the server you would otherwise be patching, not the engineering judgement.Photo: CSIRO ScienceImage 11313 The CSIRO GPU cluster at the data centre by division, CSIRO, CC BY 3.0, via Wikimedia Commons

The short answer

Serverless means you deploy a function or a container image and the provider decides when to start it, how many copies to run and when to stop. You are billed for the time your code is actually executing, measured in fractions of a second, plus a small charge per request. No instance to size, no operating system to patch, no capacity sitting idle overnight.

That model fits four shapes of work well: HTTP APIs with unpredictable traffic, scheduled jobs, event handlers reacting to a file upload or a queue message, and glue code between two systems that neither team wants to own. It fits three shapes badly: sustained high throughput, anything latency-critical at the first request, and workloads that need specific hardware or run for hours.

A quick test before you read further
  • Does your busiest hour use more than about ten times the compute of your quietest hour? Serverless is probably cheaper.
  • Does a single unit of work take longer than fifteen minutes? Serverless is probably wrong.
  • Do you need a GPU, a specific kernel module or a persistent connection held open for hours? Look at containers or virtual machines instead.

What you actually pay, with real numbers

AWS publishes Lambda’s price plainly, so this is arithmetic rather than guesswork. At the time of writing, October 2026, an x86 function costs USD 0.0000166667 per GB-second of execution plus USD 0.20 per million requests, and the monthly free allowance is one million requests and 400,000 GB-seconds. Take a modest API handler: 512 MB of memory, 200 ms per call. Each call therefore consumes 0.1 GB-seconds.

Monthly AWS Lambda bill: 512 MB function, 200 ms per request (list price, October 2026)
1 million requestsUSD 0.00 (inside the free tier)
5 million requestsUSD 2.47
10 million requestsUSD 11.80
25 million requestsUSD 39.80
Calculated from AWS Lambda’s published x86 rates of USD 0.0000166667 per GB-second and USD 0.20 per million requests, with the free allowance of 1 million requests and 400,000 GB-seconds applied. Source: AWS Lambda pricing page, accessed October 2026.

One million invocations a month land entirely inside the free allowance and cost nothing. Ten million cost about the price of a takeaway. That is the number that makes serverless attractive to small teams: a reasonably busy internal API is genuinely close to free, and you have no server to patch for it.

Two caveats before you redesign anything. First, compute is rarely the whole bill. Data transfer, the API gateway in front of the function, the database it talks to and the logs it writes often cost more than the function itself. Second, memory and duration multiply. Double the memory and you double the compute charge, although a faster function may finish in half the time and cost the same. Measuring beats guessing here, and it takes an afternoon.

Cable closet bh
Billing by the millisecond rewards short work and punishes long-running processes.Photo: Cable closet bh by Wikimedia Commons contributor, Public domain, via Wikimedia Commons

Where serverless stops being cheap

Now invert the example. Suppose your workload is steady: a container handling requests continuously, one vCPU’s worth, all month. Google Cloud’s published Cloud Run rates make the shape of the problem visible, because Google prices active and idle CPU time separately. Active CPU is USD 0.000024 per vCPU-second under instance-based billing, while idle time on a minimum instance is USD 0.0000025 per vCPU-second, roughly a tenth as much.

That ratio is the whole argument. Serverless platforms are built to charge you almost nothing when nothing is happening, and a premium when something is. Flip the duty cycle from spiky to constant and you are paying the premium rate continuously. A reserved virtual machine or a small always-on container, bought on a commitment, will usually beat it.

Serverless is not cheap compute. It is expensive compute you buy in very small amounts.

There is an operational cost on the other side of the ledger too, and it runs the other way. No instances means no operating system patching, no capacity planning and no 3 a.m. disk-full alert. For a team of three developers with no dedicated operations engineer, that saving is worth more than the compute difference on a workload of modest size. Price the engineering time, not just the invoice.

The hard limits nobody reads until they bite

AWS documents Lambda’s quotas, and three of them decide architecture. A function can run for up to 900 seconds, which is fifteen minutes, per invocation. Memory is configurable from 128 MB to 10,240 MB in one-megabyte steps, and CPU is allocated in proportion: at 1,769 MB a function has the equivalent of one vCPU. Asynchronous invocations on Lambda’s managed instances can stretch further, up to 5,400 seconds, but the fifteen-minute figure is the one to design against.

LimitValueWhat it means in practice
Maximum synchronous execution900 seconds (15 minutes)Video encoding, large report generation and bulk imports need chunking or a different service
Memory range128 MB to 10,240 MBConfigurable in 1 MB increments; CPU scales with the memory setting
One vCPU equivalent1,769 MBBelow this you get a fraction of a core, which lengthens CPU-bound work
Asynchronous and event-source invocationsUp to 5,400 seconds (90 minutes)Applies to Lambda managed instances, with exceptions for Amazon MQ and DocumentDB
AWS Lambda quotas, from the AWS Lambda developer guide, accessed October 2026.

The fifteen-minute ceiling catches people out more than anything else on that list. A nightly job that takes eleven minutes today is a time bomb, because next year’s data volume will push it past the limit and the failure will arrive as a timeout at two in the morning. Either split the work into chunks that each finish well inside the window, or run it on a container service where the clock does not stop.

Cold starts, honestly

A cold start is the delay when the platform has to create a new execution environment before your code runs. It is real, it varies by language and package size, and it matters far less than the internet suggests. For an internal tool or a webhook handler, nobody notices. For a checkout page where the first byte decides whether a customer stays, it matters a great deal.

The practical approach is to measure your own percentiles rather than trust a benchmark blog. If the ninety-ninth percentile of your first-request latency is inside your budget, stop worrying. If it is not, every major platform now sells a way around it: provisioned concurrency on Lambda, minimum instances on Cloud Run, always-ready instances on Azure Functions. Each one reintroduces a baseline cost, which is exactly the trade-off you were avoiding, so price it before you enable it.

Cheaper fixes to try first
  • Shrink the deployment package. Less to load means less to wait for.
  • Keep the runtime warm with real traffic rather than a synthetic ping, if your usage pattern allows it.
  • Move initialisation out of the request path, so the connection pool or config load happens once.
  • Choose a runtime that starts quickly for the hot paths and keep the heavier language for batch work.
The original finding aid described this photograph as: Base: Arnold Air Force Base State: Tennessee (TN) Country: United States Of America (USA) Scene Camera Operator: SMSGT. Rober
Measure your own latency percentiles before paying for provisioned capacity.Photo: A computer operator works at an IBM 4381 four-window work station in a computer room at the Arnold Engineering Development Center, where numerous mainframe and super computers are u – DPLA – b8ef28c4e9b101b7ccc7f2e5ee1a68ed by Department of Defense. American Forces Information Service. Defense Visual Information Center. 1994, Public domain, via Wikimedia Commons

Platform by platform: what the three big runtimes charge

The three major platforms have converged on a similar shape, with a free monthly grant and per-unit billing after that. The figures below come from each vendor’s own pricing page in October 2026, with the Azure numbers read from Microsoft’s public Retail Prices API for the East US region because the pricing page renders its figures dynamically.

PlatformFree monthly grantCompute priceRequest or execution priceLongest single run
AWS Lambda (x86)1M requests + 400,000 GB-sUSD 0.0000166667 per GB-secondUSD 0.20 per million requests900 s (15 min)
Azure Functions (Consumption)1M executions + 400,000 GB-sUSD 0.000016 per GB-secondUSD 0.20 per million executionsConfigurable, plan dependent
Azure Functions (Flex Consumption)250,000 executions + 100,000 GB-sUSD 0.000026 per GB-second on demandUSD 0.40 per million executionsConfigurable, plan dependent
Google Cloud Run (request billing)180,000 vCPU-s + 360,000 GiB-s + 2M requestsUSD 0.000018 per vCPU-secondIncluded in the free grant aboveConfigurable, up to hours
List prices from each vendor, October 2026. Azure Consumption and Flex Consumption figures retrieved from Microsoft’s Retail Prices API for East US; Google figures are request-based billing in tier 1 regions. Prices change, so check before you commit.

Notice how close the two headline rates are. AWS at USD 0.0000166667 per GB-second and Azure at USD 0.000016 are within a few per cent of each other, and the free grants are identical at one million calls and 400,000 GB-seconds. Price is not the deciding factor between them. The deciding factors are which ecosystem your data already lives in, which identity system your team already runs, and which runtime your developers already know.

Cloud Run deserves a mention for a different reason. It runs ordinary container images, so there is no proprietary function signature to adopt and no rewrite if you later move the same image to Kubernetes. If you want the operational simplicity of serverless without the lock-in of a vendor-specific handler, that is the most portable of the three.

A decision you can make in an afternoon

You do not need a six-week evaluation. Pick the single most annoying workload you own, the one that wakes someone up or costs more than it should, and answer five questions about it.

  1. What is the duty cycle? Hours busy per day. Under four, serverless is likely cheaper.
  2. How long does one unit of work take? Over a few minutes, be careful. Over fifteen, choose something else.
  3. What is your latency budget for a first request? If the answer is tight, plan for provisioned capacity and add its cost back in.
  4. What does it talk to? A database with a connection limit will meet a function that scales to hundreds of copies, and the database will lose. Plan for a proxy or a pool.
  5. Who maintains it today? If the honest answer is nobody, the patching you stop doing is the real win, and that argues for serverless almost regardless of the arithmetic.

For most of the small and mid-sized businesses we work with, the answer is a mixture rather than a doctrine. Scheduled jobs, webhooks and image processing go serverless. The main application stays on a container platform. Our cloud solutions team builds that split remotely, with no site visit needed, and will tell you when the honest recommendation is to leave a working system alone.

Two neighbouring decisions are worth reading before you commit. If you are weighing a container platform for the steady part of your estate, our piece on containers and Kubernetes covers when the orchestration overhead pays for itself. And whichever model you pick, the security boundary shifts with it, which is the subject of the shared responsibility model explained.

Frequently asked questions

Is serverless cheaper than a virtual machine?

For spiky workloads, usually yes, and often dramatically so. A function that handles a million requests a month inside AWS’s free allowance costs nothing, while the smallest always-on instance still has a monthly price. For steady, continuous load the position reverses, because you are paying a premium per-second rate around the clock instead of a discounted reserved rate.

What happens when my function needs more than fifteen minutes?

AWS caps a synchronous Lambda invocation at 900 seconds. The usual fixes are to split the work into chunks coordinated by a queue or a state machine, or to move that job to a container service with no execution ceiling. Asynchronous invocations on Lambda managed instances extend to 5,400 seconds, but designing against the fifteen-minute figure is safer.

Do cold starts make serverless unsuitable for customer-facing pages?

Not automatically. Measure the ninety-ninth percentile of first-request latency against your own budget. If it fails, provisioned concurrency, minimum instances or always-ready instances will fix it at the cost of a baseline charge. Many teams find that a smaller deployment package and moving initialisation out of the request path is enough.

Can I move off serverless later if it does not work out?

It depends on what you adopted. A container-based service such as Cloud Run runs ordinary images, so the same artefact runs elsewhere. Provider-specific function handlers and event bindings take more unpicking. Keeping business logic in plain functions, with a thin adapter at the edge, keeps that door open for very little extra effort.

Does serverless mean I no longer have security responsibilities?

It removes operating system patching from your list, which is a genuine reduction. Your code, your dependencies, your identity and access rules and your data handling all stay yours. Microsoft’s shared responsibility documentation keeps data, endpoints, accounts and access management in the customer column for every service model, including platform services.

Not sure whether your workload belongs on functions, containers or a plain server? We will model the three options against your real traffic and tell you which one wins. Get in touch with Eudora Technology to talk about your project.

Sources

  1. AWS Lambda pricing Amazon Web Services · accessed October 2026
  2. Lambda quotas AWS Documentation · accessed October 2026
  3. Azure Functions pricing Microsoft Azure · accessed October 2026
  4. Azure Retail Prices API Microsoft · queried October 2026
  5. Cloud Run pricing Google Cloud · accessed October 2026
  6. Kubernetes production use hits 82% in the 2025 CNCF Annual Survey Cloud Native Computing Foundation · 20 January 2026
Keep reading

Related insights