# How Notion cut embedding costs by 80%

Canonical: https://brew.new/browse/templates/email/pt1_k97zt0t83sag0fw1kd02wxs7e58e5bqh

Brand: anyscale.com
Category: newsletter

![Preview of How Notion cut embedding costs by 80%](https://cdn.brew.new/email-preview-4c0840d78a3b0e42-tracking_r57wgytakppv7xb4phbbhsv3ks8dvjgy-1788999150897.png)

## Email content

Hi Thomas,

Kicking off Ray on the Road last week, practitioners from Notion, Salesforce, Uber, and Apple took the stage at Ray Day: Seattle to share how they are running Ray in production.

A few highlights:

Notion replaced a 3-step Spark pipeline with a single Ray job on Anyscale, cutting embedding costs by 80% and improving query latency by 10x.

Salesforce summarizes documents up to 200K tokens with a P95 latency under 15 seconds using a distributed Ray actor pool running vLLM.

Uber improved GPU utilization by 20% and cut training time in half with a heterogeneous Ray cluster design.

Apple shared how Ray unifies data processing, training, and inference for foundation model workloads at scale.

Read the full recap and catch up on everything from the talks, the hands-on workshop, and what is coming next on Ray on the Road.

Read the recap

Anyscale_logo_grey_blue

Advance your AI Platform with Anyscale.

LinkedIn

X

Facebook

Anyscale, Inc., 600 Harrison St, San Francisco, CA 94107, USA

Unsubscribe

Manage preferences

[Open and remix this design](https://brew.new/browse/templates/email/pt1_k97zt0t83sag0fw1kd02wxs7e58e5bqh)

[Browse email designs](https://brew.new/browse/templates)
