Skip to content
Back to student guides
Ray ServeMLOpsModel serving3 levels103 sectionsCovers Ray Serve 2.58

The Complete Ray Serve Guide

Serve models and composed inference pipelines at scale with Ray Serve. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
16sections
32examples

This is part one of three. It covers everything you need to do real work with Ray Serve, not a teaser. By the end you can turn a Python function or a machine-learning model into a web API, run several copies of it, make one piece of your service call another, describe the whole thing in a configuration file, and read the status output when something goes wrong. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas in Serve only stick once you have watched your own service start, answer a request, and survive a crash.

You do not need to know Ray beforehand. You do need to be comfortable with basic Python (classes, functions, pip) and to have seen an HTTP request, for example with curl. If you have never built a web API at all, skim the FastAPI guide first; Serve borrows from it heavily.

What Ray Serve is, and the problem it solves

A trained model on your laptop is not a product. Somebody has to be able to send it an input over the network and get a prediction back, many times a second, without the service falling over. That job is called online inference, and Ray Serve is a library for doing it.

Serve is a scalable model-serving library built on top of Ray. It is framework-agnostic: PyTorch, TensorFlow, Keras, scikit-learn, or any ordinary Python code can sit behind it. The documentation describes it as built for end-to-end inference APIs that mix several models with business logic, rather than only "tensor in, tensor out". That phrase matters. Real prediction services rarely consist of a single model call. They clean the input, look something up, call a model, perhaps call a second model with the first one's answer, and then shape a response. Serve lets all of that live in one Python program that you can scale as a unit.

PYTHON CLASSyour model + logic
→
DEPLOYMENTdecorated with @serve.deployment
→
REPLICASseveral running copies
→
HTTP APIone address, load balanced

The diagram is the whole idea. You write a normal class, mark it as a deployment, and Serve runs as many copies of it as you ask for, spreads incoming requests across them, and exposes a single address to callers.

Two consequences of that design explain most of what follows.

Your service is ordinary Python. There is no special model format and no separate serving language. If you can call your model from a script, you can serve it. This is why people who have used heavier model servers often find Serve refreshingly small: the unit of work is a class you already know how to write.

Scaling is a setting, not a rewrite. The same class that runs as one copy on your laptop runs as ten copies on a cluster, because the copies are managed by Ray. You change a number, or you let Serve change it for you based on load. We will do the first version of that in this guide and leave the sophisticated version for later.

What people use it for:

🧠

Model APIs

Wrap a scikit-learn, PyTorch or Hugging Face model in an HTTP endpoint that several copies can share.

🔗

Pipelines of models

Chain a preprocessor, a classifier and a post-processing step as separate, separately scalable pieces.

⚡

Heavy inference

Share GPUs between replicas and handle bursts by adding replicas rather than buying a bigger machine.

💬

LLM serving

A module called Ray Serve LLM exposes OpenAI-compatible endpoints on top of engines such as vLLM.

Try it
  1. Think of a model or function you have written that someone else might want to call. Write down its input and its output as plain words.
  2. List every step between a raw request arriving and your model being called (parsing, cleaning, validating).
  3. Circle the steps that could be a separate piece. You will use that list in the composition section.

What came before Serve

To see why Serve exists, look at the options people used before it, and at where each one hurts.

The simplest option is a web framework around your model: a Flask or FastAPI app that loads the model at start-up and calls it inside a route function. This is a perfectly good way to start, and for one small model on one machine it is often all you need. The trouble appears when traffic grows. To use more CPU cores you start more worker processes, and each loads its own copy of the model. To use more machines you add a load balancer and a deployment system. If one step of your pipeline is slow and the others are fast, you cannot scale that step alone, because the whole app is one blob. If the model needs a GPU, you now have to decide by hand how many workers share it.

The second option is a dedicated model server that expects your model in a particular format and exposes a fixed interface. These are strong when your model fits the format and the request is just a tensor. They are less comfortable when your service needs arbitrary Python between the request and the model, such as calling a feature store, a vector database, or another model chosen at run time.

The third option is hand-built infrastructure: containers on Kubernetes, a queue, a custom autoscaler. This works and gives total control, but it means your team now owns a serving platform instead of a model.

Serve sits between these. You keep writing Python, which is the comfort of the first option. But replicas, load balancing, scaling and composition are handled by Ray, which removes most of the pain of the third. The honest framing, which the rest of this course will repeat, is that Serve pays off when you need several replicas, shared GPUs, multi-step composition, or LLM serving. If you have one small model on one machine, a plain FastAPI app may be the better choice, and Serve can wait until you outgrow it. The FastAPI guide covers that starting point, and you will see later that Serve can wrap a FastAPI app rather than replace it.

Reach for Serve when

  • You need several replicas of a model behind one address
  • Replicas must share or split GPUs
  • Your service is a pipeline of several models and logic steps
  • You expect load to rise and fall and want scaling to follow it

A plain web app is enough when

  • One small model runs on one machine
  • Traffic is low and steady
  • You do not need independent scaling of parts
  • You would rather not run a Ray process at all
Try it
  1. Pick the project you described in the previous section.
  2. Decide honestly which column above it belongs to today.
  3. Write one sentence describing the traffic level at which your answer would change.

The mental model: four nouns

Serve has a small vocabulary. Learn these four words and the rest of the documentation becomes readable.

A deployment is a Python class or function decorated with @serve.deployment. It is the unit of scaling: the thing you ask Serve to run several copies of. A deployment is a description, not a running process.

A replica is one running copy of a deployment. Technically it is a Ray actor, which is a long-lived Python process managed by Ray that holds state (for example, a loaded model) between requests. If you ask for three replicas, there are three such processes, each with its own copy of whatever your class loaded in its constructor.

An application is one or more deployments bound together. You create the binding by calling .bind() on a deployment, which returns an application object without running anything. The top-most deployment in an application is called the ingress: it is the one that receives outside traffic. An application has a name and a route_prefix, the URL path under which it is reachable.

A handle, called a DeploymentHandle, is how one deployment calls another at run time. When your ingress needs the preprocessing deployment, it uses a handle. Handles are what make composition possible, and we meet them in their own section.

one Ray Serve application
ProxyAccepts HTTP on port 8000 and picks a replica of the ingress
ReplicasYour class, run as Ray actors, each holding its own loaded model
ControllerTracks what should be running and restarts what dies

Behind the scenes there are two more actors that you do not write but should know exist. The controller is a single, long-lived actor that records which deployments should exist, starts and stops replicas, and drives autoscaling. The proxy is a per-node actor that accepts incoming HTTP traffic and routes each request to a replica. You will see both named in logs and status output, so it helps to recognise them early.

One vocabulary note. Ray itself is a general framework for running Python in parallel across machines, and Serve is only one library built on it. You do not need to learn Ray tasks or actors to use Serve at this level. If you are curious about the foundation, the Ray guide in this catalogue covers it.

Deployment versus replica Say it this way and you will rarely be confused: a deployment is the recipe, a replica is a dish that has been cooked from it. Changing the number of replicas changes how many dishes are on the counter, not the recipe.
Try it
  1. Without looking back, write one sentence each defining deployment, replica, application and handle.
  2. Say which of the four is a description and which is a running process.
  3. Explain in your own words why an application needs exactly one ingress.

Installing Ray Serve and checking the setup

Serve ships inside the ray Python package and is versioned together with Ray. There is no separate Serve version number. This guide was written against the Ray 2.58 line, and the documentation's current stable release at the time of writing is Ray 2.58.x. Because versions move quickly, check what pip will give you before you start:

BASH
pip index versions ray

Ray supports Linux on x86_64 and aarch64, macOS on Apple Silicon, and Windows in beta. Multi-node clusters on Windows are experimental, so if you are on Windows and want the smoothest path, use WSL2. The supported Python versions are 3.10, 3.11 and 3.12, with 3.13 in beta. Older Python versions are not supported by current Ray, so if python --version prints 3.9 or lower, install a newer interpreter first.

Always use a virtual environment. Ray pulls in many dependencies, and you do not want them tangled with the rest of your machine.

BASH
python -m venv .venv
source .venv/bin/activate
pip install -U "ray[serve]"

On Windows PowerShell the equivalent is:

POWERSHELL
py -m venv .venv
.venv\Scripts\Activate.ps1
pip install -U "ray[serve]"

The quotes around ray[serve] are not decoration. In zsh, the default shell on macOS, square brackets are special characters, and an unquoted ray[serve] fails with a "no matches found" message. Quoting it works everywhere.

The serve in square brackets is an extra: it asks pip to install Ray plus the additional packages Serve needs. There are other extras you may meet. ray[default] gives you the common set including the dashboard. ray[data,train,tune,serve] installs the machine-learning libraries together. ray[llm] adds the LLM-serving stack with vLLM and normally wants a GPU. For this guide, ray[serve] is all you need.

Verify the installation with a command that imports the library and prints its version, then check that the command-line tool is on your path:

BASH
python -c "import ray; from ray import serve; print(ray.__version__)"
serve --help

The first command should print a version such as 2.58.0. The second should list subcommands including run, build, deploy, status, config and shutdown. If the first fails with ModuleNotFoundError: No module named 'ray', your virtual environment is not active. If the second says command not found, the same cause is likely: the serve program lives in the environment's bin folder.

Behind a TLS-inspecting network If your employer or university runs a gateway that intercepts HTTPS, `pip install` and any later model download can fail with a certificate error such as `SSLCertVerificationError`. The fix is to make Python trust your organisation's root certificate, not to switch verification off. Ask your IT team for the certificate file, or export it from your device's trust store, and point the tool at it.

If you plan to run Serve inside a container, remember that Ray wants roughly 30 percent of the machine's memory as shared object store memory, and Docker's default shared-memory allowance is small. The Ray documentation recommends raising it with the --shm-size flag when you start the container. The Docker guide explains the container side of that.

Try it
  1. Create a virtual environment and install "ray[serve]".
  2. Run the version command and write down the version it prints.
  3. Run serve --help and find the six subcommands named above.

Your first deployment, step by step

Create a file called hello.py. We build it in stages so you can see what each line is for.

hello.py
from ray import serve
from starlette.requests import Request


@serve.deployment
class Hello:
    def __call__(self, request: Request) -> str:
        return "hello"


app = Hello.bind()

Read it from the top. The first line imports the serve module from Ray. The second imports Request, a class from Starlette, the small web toolkit that Serve uses to represent incoming HTTP requests. You do not install Starlette separately; it comes along with Serve.

The decorator @serve.deployment turns the class below it into a deployment. Nothing runs yet. The class has one method, __call__. In Python, defining __call__ makes an instance callable like a function, and Serve uses it as the entry point: when an HTTP request arrives, Serve calls the instance with a Request object and sends back whatever you return. Here we return the string "hello".

The last line, app = Hello.bind(), builds the application. bind does not start anything. It records that you want the Hello deployment, with whatever constructor arguments you pass (none here), as an application called app. That name matters, because the command you are about to run refers to it.

Now run it from the same folder:

BASH
serve run hello:app

The argument hello:app is an import path: the module name (the file hello.py without the extension), a colon, and the variable inside it. Serve imports hello, takes the object named app, starts a local Ray instance if one is not already running, and deploys the application. The command blocks and keeps running. By default it serves on http://127.0.0.1:8000. Leave it running and open a second terminal:

BASH
curl http://127.0.0.1:8000/

You should see hello. Back in the first terminal you will see log lines for the request. Press Ctrl+C there to stop the service.

Congratulations: you have run an online inference service. It does not contain a model yet, but the machinery is complete. A request went through the proxy to a replica of your deployment and a response came back.

Reading the logs The lines that mention a replica and the route you called tell you which actor handled the request. When something later does not work, the first thing to do is read these lines, because they usually name the deployment and the exception.

You can also run an application from inside Python rather than from the command line:

PYTHON
from ray import serve

handle = serve.run(app, name="hello", route_prefix="/")

serve.run deploys the application, gives it a name and a route prefix, and returns a handle to the ingress deployment. It is handy in notebooks and tests. For anything you keep, the command line and, later, the configuration file are the better habit.

Try it
  1. Create hello.py as shown and run serve run hello:app.
  2. From a second terminal, call it with curl and confirm you get hello.
  3. Change the returned string, stop the service with Ctrl+C, run it again, and confirm the new text appears.

Handling requests and returning JSON

A real service accepts input. In the first example the method received a Request and ignored it. Now we use it.

The Request object is a Starlette request, so it offers the usual accessors. request.query_params holds the values after the question mark in the URL. await request.json() reads a JSON body. Because reading the body is asynchronous, a method that does so must be declared async def.

greeter.py
from ray import serve
from starlette.requests import Request


@serve.deployment
class Greeter:
    def __init__(self, greeting: str = "Hello"):
        self.greeting = greeting

    async def __call__(self, request: Request) -> dict:
        body = await request.json()
        name = body.get("name", "stranger")
        return {"message": f"{self.greeting}, {name}!"}


app = Greeter.bind("Salam")

Notice three things. First, the constructor __init__ runs once per replica, when the replica starts. This is the right place to load a model, open a connection, or read a file, because you pay the cost once and every later request reuses the result. Second, the constructor takes an argument, and Greeter.bind("Salam") supplies it. Whatever you pass to bind goes to the constructor. Third, the method returns a dictionary, and Serve converts it to a JSON response.

Run it with serve run greeter:app and call it:

BASH
curl -X POST http://127.0.0.1:8000/ \
  -H "Content-Type: application/json" \
  -d '{"name": "Aya"}'

The reply is {"message": "Salam, Aya!"}.

Why do we care about the difference between __init__ and __call__? Because it is the single most important performance habit in model serving. Loading a model can take seconds or minutes. If you load it inside __call__, every request pays that cost. If you load it in __init__, only the start-up pays it. A beginner who puts pipeline("summarization") inside the request method will see a service that works but takes many seconds per call, and will blame the model. The model is fine; the loading is in the wrong place.

The other thing to understand is sync versus async. A method declared def runs in a thread pool on the replica, so slow blocking work in it does not freeze the replica. A method declared async def runs on the replica's event loop, which is faster for waiting on network calls but has a rule: never put a long blocking call (a heavy model forward pass, time.sleep) directly in it, because that stalls everything else the replica is doing. For a beginner, the safe rule is this: use async def when you need to await something such as request.json(), and keep the heavy computation short or move it into a plain def method that you call.

Blocking inside async A long, blocking call inside an `async def` method freezes that replica's event loop. Requests queue up behind it and look like timeouts. If you see a replica that answers one request at a time despite low CPU use, check for blocking code in async methods first.
Try it
  1. Create greeter.py and run it.
  2. Send two requests with different names and confirm the greeting word stays the same.
  3. Change the argument passed to bind, restart, and confirm the greeting changes.

Serving a real model with FastAPI

Hand-parsing JSON gets tedious fast. Serve integrates with FastAPI, which gives you routing, input validation and automatic documentation, and the integration is one decorator. If you know FastAPI, you already know most of this section; the FastAPI guide is the place to learn it properly.

The pattern is to create a FastAPI app object, decorate your deployment class with @serve.ingress(api), and declare routes on the class with the app's decorators. The class becomes the ASGI application (ASGI is the interface standard that Python async web apps share).

text_ml.py
from fastapi import FastAPI
from ray import serve

api = FastAPI()


@serve.deployment
@serve.ingress(api)
class Summarizer:
    def __init__(self):
        from transformers import pipeline

        self.pipe = pipeline("summarization")

    @api.post("/")
    def summarize(self, text: str) -> str:
        return self.pipe(text)[0]["summary_text"]


app = Summarizer.bind()

To run this you also need the libraries it imports:

BASH
pip install transformers torch
serve run text_ml:app

Look at the order of decorators. @serve.deployment is on the outside and @serve.ingress(api) sits just above the class. The route is declared with @api.post("/"), exactly as in a normal FastAPI app. The text: str parameter is read by FastAPI from the request, and the return annotation tells FastAPI how to shape the response.

The import of transformers is inside __init__ instead of at the top of the file. This is a deliberate choice that the Serve documentation recommends: a delayed import means the heavy library is imported on the replica, where the model runs, rather than in the process that merely builds the application object. It also lets different deployments in one application use different dependency versions later on.

Test it. Because text is a plain str parameter on a POST route, FastAPI expects it as a query parameter:

BASH
curl -X POST "http://127.0.0.1:8000/?text=Ray%20Serve%20lets%20you%20turn%20Python%20classes%20into%20scalable%20web%20services."

The first run downloads the default summarization model, which can take a few minutes and needs internet access, so be patient and watch the logs. After that, requests are fast. Since FastAPI is in control, you also get interactive documentation for free: open http://127.0.0.1:8000/docs in a browser to see and try your routes.

Which style should you use, the raw Request or FastAPI? For anything beyond a quick experiment, FastAPI. You get validation (a request with a missing field is rejected with a clear error before your code runs), documentation, and a style that other engineers recognise. The raw Request approach is fine for tiny services or when you need full control of the body.

Swap the model, keep the shape Nothing about Serve is specific to Transformers. Replace the pipeline with a scikit-learn model loaded using `joblib.load` in `__init__`, and call `predict` in the route. The structure (load once in the constructor, predict in the method) is identical.
Try it
  1. Install the two extra packages and run text_ml.py.
  2. Call it with curl and then open the /docs page.
  3. Send a request with no text parameter and read the validation error FastAPI returns.

Replicas and resources: running more than one copy

So far each deployment has had a single replica. The reason to use Serve at all is that you can ask for more, and the way to do it is an argument to the decorator.

scaled.py
from ray import serve
from starlette.requests import Request


@serve.deployment(
    num_replicas=2,
    ray_actor_options={"num_cpus": 1},
)
class Model:
    def __init__(self):
        import os

        self.pid = os.getpid()

    async def __call__(self, request: Request) -> dict:
        return {"served_by_process": self.pid}


app = Model.bind()

num_replicas=2 tells Serve to run two copies. ray_actor_options describes what each replica needs from the machine. Here every replica reserves one CPU. Ray treats these numbers as logical resources: they are a promise used for scheduling, not a hard limit enforced on the process. If a machine has eight CPUs and each replica asks for one, Ray will place up to eight such replicas there. If you ask for more than the cluster has, the extra replicas wait, and the status output will tell you they are waiting for resources.

Run it and call it several times:

BASH
serve run scaled:app
# in another terminal
for i in 1 2 3 4 5 6; do curl -s http://127.0.0.1:8000/; echo; done

You will see two different process ids alternating or mixing in the answers. That is load balancing, made visible: the proxy spreads the requests across both replicas. Each replica ran its own __init__, so each has its own pid, and, with a real model, its own copy of the model weights in memory. This is the first thing to budget for. Two replicas of a 2 GB model need about 4 GB of memory, not 2.

For models that need a GPU, you ask in the same place:

PYTHON
@serve.deployment(
    num_replicas=2,
    ray_actor_options={"num_cpus": 1, "num_gpus": 0.5},
)
class GpuModel:
    ...

Fractional GPUs are allowed. Two replicas that each request half a GPU share one physical card, which can save a lot of money when a model is small enough that a whole GPU would sit mostly idle. Note the word "logical" again: Ray will not stop one replica from using more than half of the card's memory. Fractions are a scheduling agreement, and keeping each model inside its share is your responsibility.

There is a second knob you will see in the documentation and should recognise: max_ongoing_requests. This limits how many requests a single replica handles at once. Its default is 5 in current versions (it was 100 before Ray 2.32, and was named max_concurrent_queries, a name you may still see in older tutorials). If every replica is already handling its maximum, new requests wait in a queue. For a model that processes one request at a time, a small number is sensible, and it is a safety valve that stops one replica from accepting more work than it can handle.

A deployment also has a name. By default it is the class name. You can set it explicitly with @serve.deployment(name="summarizer"), which makes status output easier to read when you have several deployments.

Old names in old tutorials Tutorials from before 2024 use `max_concurrent_queries` and `target_num_ongoing_requests_per_replica`. These were renamed to `max_ongoing_requests` and `target_ongoing_requests`, and the defaults changed in Ray 2.32. If code you copy from the internet fails on an unknown option, check the date of the post before you debug anything else.
Try it
  1. Run scaled.py and send ten requests. Count how many different process ids you see.
  2. Change num_replicas to 4 and repeat. Count again.
  3. Ask for more CPUs per replica than your machine has and read what happens in the logs.

Calling one deployment from another

Real services have parts. Suppose a request carries messy text, and the model wants it cleaned first. You could do both in one class, but then cleaning and modelling scale together even if one of them is trivial. Serve lets you split them into two deployments and connect them.

The connector is the handle. When you pass one bound deployment into another's bind call, Serve gives the receiving class a DeploymentHandle for it in its constructor. You call the handle's method with .remote(...), and you get back a DeploymentResponse, which you can await.

pipeline.py
from ray import serve
from starlette.requests import Request


@serve.deployment
class Preprocessor:
    def __call__(self, text: str) -> str:
        return text.strip().lower()


@serve.deployment
class Model:
    def __init__(self, pre):
        self.pre = pre

    async def __call__(self, request: Request) -> dict:
        text = (await request.json())["text"]
        clean = await self.pre.remote(text)
        return {"out": clean}


app = Model.bind(Preprocessor.bind())

Read the last line carefully, because it is where composition happens. Preprocessor.bind() creates a bound preprocessor. Model.bind(...) receives it as an argument, so Model's constructor gets pre, which is a handle to the preprocessor, not the preprocessor itself. In __call__, self.pre.remote(text) sends the text to a preprocessor replica and returns immediately with a response object. Awaiting it produces the cleaned string.

The top of the chain, Model, is the ingress. It is the only deployment reachable from outside. Preprocessor has no route of its own; it is an internal step. This is a useful property: you can scale the two parts independently, give them different resources, and give the preprocessor no GPU while the model has one.

PYTHON
@serve.deployment(num_replicas=1, ray_actor_options={"num_cpus": 0.5})
class Preprocessor:
    ...


@serve.deployment(num_replicas=3, ray_actor_options={"num_cpus": 1, "num_gpus": 1})
class Model:
    ...
CLIENTPOST /
→
MODELingress, 3 replicas
→
PREPROCESSORreached by a handle
→
RESPONSEJSON back to client

Why await? The DeploymentResponse is a future: a placeholder for a result that is still being computed on another replica. Awaiting it lets the replica serve other requests while it waits. From ordinary synchronous code, you can call .result() on the response instead, which blocks until the answer is ready. Inside an async def method, always prefer await.

If you have read older tutorials, you may have seen something called a deployment graph with a DAGDriver, and handle types named RayServeHandle. Those belong to older Ray versions and are replaced by the .bind() composition and DeploymentHandle shown here. If you meet them, translate rather than copy.

Composition is also where cross-links to the rest of this catalogue become natural. A model-serving pipeline often reads features from a store, and logs its predictions for monitoring. When you reach for those, look at the Feast guide for features and the Evidently guide for monitoring.

Try it
  1. Create pipeline.py and run it.
  2. Send {"text": " HeLLo "} as a JSON body and confirm the cleaned output.
  3. Add a third deployment that counts words and chain it after the preprocessor.

The command line and the config file

Running serve run in a terminal is perfect for development. It is the wrong tool for something that must keep running. For that, Serve has a declarative workflow: describe what you want in a YAML file, and tell a running Ray cluster to make it so.

The command-line tool has a handful of subcommands, and it is worth knowing each by its purpose:

Command What it does
serve run Deploy an application and stay in the foreground; Ctrl+C stops it
serve build Generate a config file from an import path
serve deploy Send a config file to a running Ray cluster
serve status Show the state of the applications and deployments
serve config Print the config the cluster is currently running
serve shutdown Remove all applications from the cluster
serve start Start Serve on a cluster without deploying an application

The full workflow on one machine is as follows. First, generate a config from your import path:

BASH
serve build text_ml:app -o serve_config.yaml

This writes a YAML file that describes your application. A trimmed example looks like this:

serve_config.yaml
proxy_location: EveryNode
http_options:
  host: 0.0.0.0
  port: 8000
applications:
  - name: text_app
    route_prefix: /
    import_path: text_ml:app
    runtime_env:
      pip: ["transformers", "torch"]
    deployments:
      - name: Summarizer
        num_replicas: 2
        max_ongoing_requests: 10
        ray_actor_options:
          num_cpus: 1

Read it as a description of the same things you set in Python. applications is a list, so one cluster can host several applications, each with its own name and route_prefix. import_path is the same module:variable string you passed to serve run. runtime_env says what packages the replicas need. deployments lets you override per-deployment settings such as the number of replicas, so you can change scaling without touching the Python code.

The point of the file is that it can be committed to Git, reviewed, and applied repeatedly. The Python code defines what the service does; the YAML defines how it runs. That separation is the reason most teams use the config file in production.

Second, start a Ray cluster. On a single machine, one head node is enough:

BASH
ray start --head

Third, deploy the config and check the result:

BASH
serve deploy serve_config.yaml
serve status

Here is something that surprises newcomers. When serve deploy prints a success message, it means only that the cluster received the request. The deployment carries on in the background: replicas must start, models must load. So after deploying, run serve status repeatedly until the application reports that it is running. A model that takes two minutes to load will show a deploying state for two minutes, and that is normal.

To talk to a cluster that is not on your machine, the deploy command takes an address for the Ray dashboard, which listens on port 8265:

BASH
serve deploy serve_config.yaml -a http://<head-node-address>:8265

When you are done, clear the cluster:

BASH
serve shutdown
ray stop

The import path deserves a warning, because it is where many first deployments fail. The documentation states that the import path (for example text_ml:app) must be importable by Serve at runtime. On your laptop, the folder you run from is on Python's search path, so it works. On a cluster, the replicas are different processes, possibly on different machines, and they do not automatically have your file. The fix is to ship your code with the application: through a working_dir in the runtime_env, through a container image that contains it, or through a remote location the cluster can download from. For learning on one machine you will not hit this; remember it for the day you move to a real cluster.

The dashboard is not public Port 8265 serves the Ray dashboard and the job interface, and anyone who can reach it can run code on your cluster. Keep it on localhost or inside a trusted private network. Never open it to the internet to make a `serve deploy` convenient.
Try it
  1. Run serve build on the pipeline from the previous section and open the YAML it writes.
  2. Start Ray with ray start --head and deploy the file with serve deploy.
  3. Run serve status until the application is running, call it, then clean up with serve shutdown and ray stop.

Scaling with load, and a first look at batching

Setting num_replicas to a fixed number is fine until traffic changes. At three in the morning you pay for replicas that do nothing, and at lunchtime the same number may be too few. Autoscaling lets Serve adjust the replica count to the load.

You enable it by giving the deployment an autoscaling_config instead of a fixed num_replicas:

PYTHON
@serve.deployment(
    autoscaling_config={
        "min_replicas": 1,
        "max_replicas": 5,
        "target_ongoing_requests": 2,
    },
    max_ongoing_requests=5,
)
class Model:
    ...

Here is the idea, in plain words. Serve measures how many requests each replica is working on at once, on average. You tell it the target you are comfortable with, target_ongoing_requests. If the average is above the target, Serve adds replicas, up to max_replicas. If it is well below, Serve removes replicas, down to min_replicas. The upper and lower bounds are your guard rails: max_replicas stops a traffic spike from creating unlimited cost, and min_replicas keeps some capacity warm.

A few details prevent confusion. The default min_replicas is 1, and the default max_replicas is also 1, which means that if you add an autoscaling config but forget max_replicas, nothing will ever scale up. The default target_ongoing_requests is 2. Scaling up waits for a delay (30 seconds by default) before acting, and scaling down waits much longer (600 seconds by default), so a short quiet moment does not tear down capacity you will need a minute later. That slow scale-down is deliberate, not a bug. Finally, do not set both num_replicas and autoscaling_config on the same deployment; choose one.

You can set min_replicas to 0 to save money when the service is idle, but then the first request after a quiet period must wait for a replica to start and load its model. For a model that loads in seconds that is acceptable; for one that loads in minutes it will feel like an outage. Start with a minimum of 1 and reduce it only when you have measured the cold start.

Autoscaling changes the number of replicas. Batching changes how much work each replica does per call, and for neural network models it is often the bigger win. GPUs and vectorised libraries process a group of inputs together almost as cheaply as one. Serve offers @serve.batch to collect individual requests into a group automatically.

batched.py
from ray import serve
from starlette.requests import Request


@serve.deployment
class Batched:
    @serve.batch(max_batch_size=8, batch_wait_timeout_s=0.1)
    async def handle(self, items: list[int]) -> list[int]:
        return [i * 2 for i in items]

    async def __call__(self, request: Request) -> int:
        return await self.handle(int(request.query_params["n"]))


app = Batched.bind()

Notice the contract. The decorated method is async, it receives a list of inputs, and it must return a list of the same length, in the same order. The caller, however, passes one item and gets one result: __call__ sends a single integer and receives a single integer. Serve does the grouping behind the scenes. It waits up to batch_wait_timeout_s seconds (0.1 here), or until max_batch_size items have arrived (8 here), whichever comes first, then runs the method once on the whole group and hands each caller its own result.

The trade-off is latency against throughput. A longer wait gives bigger batches and better hardware use, but each request may wait a little longer. The documentation advises choosing a wait time below your latency budget minus the inference time, and using powers of two for the batch size. The defaults are a batch size of 10 and a wait of 0.01 seconds. You do not need to batch a service until you measure that you need it, but knowing the tool exists will save you from inventing your own.

Scale on measurement Do not guess replica counts. Send realistic load at the service, watch `serve status` and the logs, and then set the target. A number chosen from a measurement is almost always better than a number chosen from instinct.
Try it
  1. Run batched.py and send one request with curl "http://127.0.0.1:8000/?n=4". Note the delay of about a tenth of a second.
  2. Send eight requests at the same moment from a shell loop with & and confirm each caller gets its own doubled number.
  3. Reduce batch_wait_timeout_s to 0.01 and compare the single-request delay.

Configuration you will meet early

Beyond replicas and resources, a few settings come up in nearly every real project. You do not need to memorise them. You need to know where to look.

Per-deployment options are set either in the decorator or, for an existing application, in the config file's deployments list. The decorator is the default; the YAML overrides it. If a value in the file disagrees with the code, the file wins, which is intentional: it lets an operator change scaling without a code change.

Runtime environments (runtime_env) describe the Python packages and files a replica needs. In the config file you saw pip: ["transformers", "torch"]. You can also set one per deployment with ray_actor_options={"runtime_env": {"pip": [...]}}, which lets two deployments in one application use different versions of a library. A caution from the documentation: avoid packages that compile from source at start-up, because that burns resources and can destabilise a cluster. For production, prefer to bake dependencies into a container image, as described in the Docker guide.

Logging is controlled with a LoggingConfig per deployment, available from Ray 2.9. It sets the log level, a JSON output format and whether to log each request. Logs live on the node under /tmp/ray/session_latest/logs/serve/, and the Python logger is called ray.serve. When a deployment fails to start, the answer is almost always in the replica logs there.

The proxy can run on every node or only on the head node, controlled by proxy_location; the config file generated by serve build shows EveryNode. On one laptop this changes nothing.

Environment variables matter occasionally. RAY_ADDRESS and RAY_DASHBOARD_ADDRESS tell the command-line tools which cluster to talk to, so you do not have to pass -a every time.

Authentication is not something Serve does for you. The HTTP endpoints have no built-in login. If your service will be reachable by anyone other than you, put an API gateway or ingress with authentication in front, or add a check inside your FastAPI app. Secrets such as API keys should come from the environment of the replica or from a secret manager and be read in __init__; never write them into the config file you commit. The Kubernetes guide explains the secret mechanisms used on the cluster side.

Try it
  1. Edit the config you built earlier to set num_replicas to 3 and redeploy it with serve deploy.
  2. Run serve config and confirm the new value is what the cluster is running.
  3. Find the log directory named above on your machine and open the newest file.

Common errors and how to read them

Most Serve problems fall into a few families. The skill is to look in the right place first. Two places answer almost everything: serve status, and the replica logs in /tmp/ray/session_latest/logs/serve/.

The status stays stuck in a deploying or unhealthy state. This is the most common complaint. It means the replicas could not start. The reasons are usually one of three. The constructor raised an exception, for example a missing file or a failed model download. The cluster does not have the resources you asked for, so the replica is waiting; the status will say it is waiting for resources, and the fix is to ask for less or add capacity. Or the options passed in ray_actor_options are wrong. Open the replica log: the Python traceback is there, and it names your file and line.

ModuleNotFoundError for your own module, or a failure to import the application. The import path is not importable where the replica runs. On one machine, check that you run serve run or serve build from the folder containing the file, and that the module name matches the file name without .py. On a cluster, ship the code with working_dir or a container image, as described earlier.

Cannot connect to Ray, or a failure to reach the head node. You are running a command that needs a cluster, and none is running at the address. Start one with ray start --head, or correct RAY_ADDRESS. Note that serve run starts a local instance for you, but serve deploy and serve status talk to an existing one.

Connection refused on port 8265. The dashboard is not reachable from where you are. On the same machine, check Ray is running. From another machine, you will need the head node's address and a network path to it, and you should think carefully about the security warning above before opening one.

A replica dies with an actor-died error. The process was killed, very often by the operating system for using too much memory. Two replicas of a large model use twice the memory of one. Reduce the replica count or batch size, or request more memory for the replica through ray_actor_options.

Requests are slow or time out. First ask whether the model is being loaded inside __call__ instead of __init__. Then ask whether a blocking call sits inside an async def. Then ask whether max_ongoing_requests is so small, relative to traffic, that requests are queuing. Only then look at the model itself.

Complaints about both num_replicas and autoscaling_config. You set both on one deployment. Remove one.

Validation or import errors after upgrading Ray. Recent Ray releases expect Pydantic version 2. Pydantic 1 has been deprecated, and the installation documentation announced that support would be dropped in release 2.56. If a library in your environment pins the old version, you will see import errors; upgrade that library or use a clean environment.

A certificate error during model or package download. If you are behind a gateway that inspects HTTPS, make Python trust the organisation's certificate as described in the installation section, rather than disabling verification.

A debugging order that works First, read `serve status`. Second, open the newest replica log. Third, reproduce the problem outside Serve by importing your class in a plain Python shell and calling the method directly. If it fails there, the bug is in your code, not in Serve.
Try it
  1. Add raise ValueError("boom") to a constructor and run the service. Read the traceback in the log and in the terminal.
  2. Ask for "num_gpus": 1 on a laptop with no GPU and read what serve status reports.
  3. Fix both and confirm the application becomes healthy.

Putting it all together

Here is one small end-to-end project that uses most of the guide. We will build a sentiment service with two deployments, validate its input with FastAPI, scale the model, write a config, deploy it to a local Ray cluster, and call it. The model is a small scikit-learn text classifier trained inside the file, so you need no download beyond the library itself.

First install the extras:

BASH
pip install scikit-learn

Then create sentiment.py:

sentiment.py
from fastapi import FastAPI
from pydantic import BaseModel
from ray import serve

api = FastAPI()


class Review(BaseModel):
    text: str


@serve.deployment
class Cleaner:
    def __call__(self, text: str) -> str:
        return " ".join(text.lower().split())


@serve.deployment(num_replicas=2, ray_actor_options={"num_cpus": 1})
class Classifier:
    def __init__(self):
        from sklearn.feature_extraction.text import TfidfVectorizer
        from sklearn.linear_model import LogisticRegression
        from sklearn.pipeline import make_pipeline

        texts = [
            "great product, works well",
            "excellent quality and fast delivery",
            "love it, very happy",
            "terrible, it broke in a day",
            "awful experience, want a refund",
            "bad quality and slow support",
        ]
        labels = ["positive", "positive", "positive", "negative", "negative", "negative"]
        self.model = make_pipeline(TfidfVectorizer(), LogisticRegression())
        self.model.fit(texts, labels)

    def predict(self, text: str) -> str:
        return self.model.predict([text])[0]


@serve.deployment
@serve.ingress(api)
class Api:
    def __init__(self, cleaner, classifier):
        self.cleaner = cleaner
        self.classifier = classifier

    @api.post("/sentiment")
    async def sentiment(self, review: Review) -> dict:
        clean = await self.cleaner.remote(review.text)
        label = await self.classifier.predict.remote(clean)
        return {"text": clean, "label": label}


app = Api.bind(Cleaner.bind(), Classifier.bind())

Take it apart. Review is a Pydantic model: FastAPI uses it to validate the JSON body, so a request without text is rejected before your code runs. Cleaner is a trivial preprocessing step. Classifier trains a tiny model in __init__, once per replica, and exposes a predict method. Because it has two replicas, two copies of the model exist. Api is the ingress: it declares the route with FastAPI, receives handles to the other two in its constructor, and chains the calls. The last line builds the application by binding the three deployments together.

One new detail: self.classifier.predict.remote(clean) calls the predict method by name through the handle. When a deployment has a method other than __call__ that you want to reach, you call it as an attribute of the handle followed by .remote.

Run it first the quick way, to check that it works:

BASH
serve run sentiment:app
BASH
curl -X POST http://127.0.0.1:8000/sentiment \
  -H "Content-Type: application/json" \
  -d '{"text": "  Great PRODUCT, very happy  "}'

The reply is a JSON object with the cleaned text and a label, which should be positive for this input. Send an obviously negative sentence to see the other label. Send {} and read the 422 validation error from FastAPI. Stop the service with Ctrl+C.

Now do it the production way. Build the config, start a cluster, deploy and watch:

BASH
serve build sentiment:app -o sentiment_config.yaml
ray start --head
serve deploy sentiment_config.yaml
serve status

Edit sentiment_config.yaml to raise the classifier to three replicas, redeploy, and watch serve status pick up the change without you touching the Python file. Call the service again to confirm it still answers. When you are finished:

BASH
serve shutdown
ray stop

Finally, store the work properly. Put sentiment.py, the YAML and a requirements.txt in a Git repository. That repository is a complete, reviewable description of a service. When you are ready to run it for real users, the next step is a container image and a Kubernetes manifest, and the Docker guide and Kubernetes guide are the natural continuation.

Try it
  1. Build and run the sentiment service exactly as shown and test one positive and one negative sentence.
  2. Deploy it through the config file and change the classifier replicas from the YAML only.
  3. Add a field to the Review model (for example language) and observe how FastAPI validates it.

What you can now do, and what comes next

You can now take a Python class and turn it into a service: decorate it, bind it, run it with serve run, and call it with curl. You understand the four nouns, deployment, replica, application and handle, and you can explain why a model belongs in __init__. You can run several replicas, give them CPUs or a fraction of a GPU, wire two deployments together, describe the whole application in a YAML file, deploy it to a local Ray cluster, and read serve status and the logs when something misbehaves. You know the most common errors by family, and you know to check the date on a tutorial because several option names changed in 2024.

That is enough to ship a first internal service. If you work for an employer in the Gulf or in Egypt, one practical note applies before you deploy anything with real user data: decide where the cluster runs. Data-residency rules at many organisations require that personal data stay inside the country or region, and your replicas, logs and any model downloads all count as places where data may travel. Pick your cloud region on purpose.

The mid-level guide goes under the hood: how the controller, proxy and replicas cooperate, how requests are routed and queued, deeper autoscaling settings, model multiplexing, streaming responses, testing a deployment, and running Serve on Kubernetes with KubeRay. The senior guide covers failure modes, security, multi-tenancy, cost and upgrades, and when to choose something else.

Where to go next in this catalogue depends on what you are serving. For language models, the vLLM guide covers the engine that Ray Serve LLM can drive. For comparison with other serving tools, see BentoML, KServe and Triton. For metrics and dashboards on a running service, see Prometheus and Grafana.

Try it
  1. Write a one-paragraph description of a service you would build with Serve, naming its deployments and which ones need a GPU.
  2. List three settings from this guide you would put in the config file rather than the code.
  3. Decide where you would run it and why, with data residency in mind.

Sources