本页包含使用 transformWithState 该操作符的自定义有状态流应用的代码示例。 Databricks 建议对常见操作(如聚合和连接)使用内置有状态方法。
请参阅使用 transformWithState 构建自定义有状态应用程序。
注释
Python 支持基于 transformWithState 行的 API(在微批处理模式和实时模式下可用)和基于 transformWithStateInPandas Pandas 的运算符。 以下示例提供在 transformWithStateInPandas Python 和 transformWithState Scala 中的代码。
注释
本页可运行的示例是在专用 main.stateful_examples 模式中创建表,这样它们可以运行而不影响你现有的数据。 如果你没有权限在 main 目录中创建模式,请将示例中的目录和模式更改为你可以创建表的位置。
要求
transformWithState运算符和相关 API 和类具有以下要求:
- 在 Databricks Runtime 16.2 及更高版本中可用。
- Databricks Runtime 16.3 及更高版本支持 Python(
transformWithStateInPandas以及基于行的transformWithState)的标准访问模式,Databricks Runtime 17.3 及更高版本支持 Scala(transformWithState)的标准访问模式。 - RocksDB 是 Databricks Runtime 17.3 及更高版本中的默认状态存储提供程序。 对于低于 17.3 的 Databricks Runtime 版本,必须配置 RocksDB 状态存储提供程序。 Databricks 建议在计算配置过程中启用 RocksDB。
注释
在低于 17.3 的 Databricks Runtime 版本上,通过运行以下命令为当前会话启用 RocksDB 状态存储提供程序:
spark.conf.set("spark.sql.streaming.stateStore.providerClass", "org.apache.spark.sql.execution.streaming.state.RocksDBStateStoreProvider")
缓慢变化维度 (SCD) 类型 1
以下代码是一个使用 transformWithStateSCD 类型 1 实现的示例。 SCD 类型 1 仅跟踪给定字段的最新值。
注释
可以使用流式处理表,并使用 AUTO CDC ... INTO Delta Lake 支持的表实现 SCD 类型 1 或类型 2。 此示例在状态存储中实现 SCD 类型 1,为准实时应用程序提供较低的延迟。
Python
# Import the necessary libraries
import pandas as pd
from pyspark.sql.streaming import StatefulProcessor, StatefulProcessorHandle
from pyspark.sql.types import StructType, StructField, LongType, StringType
from typing import Iterator
# Set the state store provider to RocksDB
spark.conf.set("spark.sql.streaming.stateStore.providerClass", "org.apache.spark.sql.execution.streaming.state.RocksDBStateStoreProvider")
# Define the output schema for the streaming query
output_schema = StructType([
StructField("user", StringType(), True),
StructField("time", LongType(), True),
StructField("location", StringType(), True)
])
# Define a custom StatefulProcessor for slowly changing dimension type 1 (SCD1) operations
class SCDType1StatefulProcessor(StatefulProcessor):
def init(self, handle: StatefulProcessorHandle) -> None:
self.handle = handle
# Define the schema for the state value
value_state_schema = StructType([
StructField("user", StringType(), True),
StructField("time", LongType(), True),
StructField("location", StringType(), True)
])
# Initialize the state to store the latest location for each user
self.latest_location = handle.getValueState("latestLocation", value_state_schema)
def handleInputRows(self, key, rows, timerValues) -> Iterator[pd.DataFrame]:
# Find the row with the maximum time value
max_row = None
max_time = float('-inf')
for pdf in rows:
for _, pd_row in pdf.iterrows():
time_value = pd_row["time"]
if time_value > max_time:
max_time = time_value
max_row = tuple(pd_row)
# Check whether state exists and update if necessary
exists = self.latest_location.exists()
if not exists or max_row[1] > self.latest_location.get()[1]:
# Update the state with the new max row
self.latest_location.update(max_row)
# Yield the updated row
yield pd.DataFrame(
{"user": (max_row[0],), "time": (max_row[1],), "location": (max_row[2],)}
)
# Yield an empty DataFrame if no update is needed
yield pd.DataFrame()
def close(self) -> None:
# No cleanup needed
pass
import uuid
# Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
# Seed a small Delta table to use as the streaming source
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.scd1_source")
spark.createDataFrame(
[("u1", 1, "NYC"), ("u1", 3, "SF"), ("u1", 2, "LA"), ("u2", 5, "London")],
"user string, time long, location string",
).write.saveAsTable("main.stateful_examples.scd1_source")
df = spark.readStream.table("main.stateful_examples.scd1_source")
# Apply the stateful transformation to the input DataFrame
q = (
df.groupBy("user")
.transformWithStateInPandas(
statefulProcessor=SCDType1StatefulProcessor(),
outputStructType=output_schema,
outputMode="Update",
timeMode="None",
)
.writeStream.format("memory")
.queryName("scd1_output")
.option("checkpointLocation", f"/tmp/checkpoint_{uuid.uuid4()}")
.trigger(availableNow=True)
.start()
)
q.awaitTermination()
# Each user keeps only its latest location by time: u1 -> SF (time 3), u2 -> London (time 5)
display(spark.sql("SELECT user, time, location FROM scd1_output ORDER BY user"))
Scala(编程语言)
import org.apache.spark.sql.streaming._
// Define a case class to represent user location data
case class UserLocation(
user: String,
time: Long,
location: String)
// Define a stateful processor for slowly changing dimension type 1 (SCD1) operations
class SCDType1StatefulProcessor extends StatefulProcessor[String, UserLocation, UserLocation] {
import org.apache.spark.sql.{Encoders}
// Transient value state to store the latest location for each user
@transient private var _latestLocation: ValueState[UserLocation] = _
private val userLocationEncoder = Encoders.product[UserLocation]
// Initialize the state store
override def init(
outputMode: OutputMode,
timeMode: TimeMode): Unit = {
// Create a value state named "locationState" using UserLocation encoder
// TTLConfig.NONE means the state has no expiration
_latestLocation = getHandle.getValueState[UserLocation]("locationState",
userLocationEncoder, TTLConfig.NONE)
}
// Process input rows and update state
override def handleInputRows(
key: String,
inputRows: Iterator[UserLocation],
timerValues: TimerValues): Iterator[UserLocation] = {
// Find the location with the maximum timestamp from input rows
val maxNewLocation = inputRows.maxBy(_.time)
// Update state and emit output if:
// 1. No previous state exists, or
// 2. New location has a more recent timestamp than the stored one
if (_latestLocation.getOption().isEmpty || maxNewLocation.time > _latestLocation.get().time) {
_latestLocation.update(maxNewLocation)
Iterator.single(maxNewLocation) // Emit the updated location
} else {
Iterator.empty // No update needed, emit nothing
}
}
}
import spark.implicits._
import java.util.UUID
// Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
// Seed a small Delta table to use as the streaming source
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.scd1_source_scala")
Seq(
UserLocation("u1", 1L, "NYC"),
UserLocation("u1", 3L, "SF"),
UserLocation("u1", 2L, "LA"),
UserLocation("u2", 5L, "London")
).toDF().write.saveAsTable("main.stateful_examples.scd1_source_scala")
val q = spark.readStream
.table("main.stateful_examples.scd1_source_scala")
.as[UserLocation]
.groupByKey(_.user)
.transformWithState(
new SCDType1StatefulProcessor(),
TimeMode.None(),
OutputMode.Update()
)
.writeStream
.format("memory")
.queryName("scd1_output_scala")
.option("checkpointLocation", s"/tmp/checkpoint_${UUID.randomUUID()}")
.trigger(Trigger.AvailableNow())
.start()
q.awaitTermination()
// Each user keeps only its latest location by time: u1 -> SF (time 3), u2 -> London (time 5)
spark.sql("SELECT user, time, location FROM scd1_output_scala ORDER BY user").show()
缓慢变化维度 (SCD) 类型 2
以下笔记本包含一个示例,展示如何在 Python 或 Scala 中使用 transformWithState 实现 SCD 类型 2。
SCD 类型 2 Python
SCD 类型 2 Scala
故障时间检测器
transformWithState 可实现计时器,以支持你根据已用时间执行操作,即使在微批处理中未处理给定键的任何记录也是如此。
以下示例展示停机检测器模式的实现。 每次看到给定键的新值时,都会更新 lastSeen 状态值,清除任何现有计时器,并为将来重置计时器。
计时器过期时,应用程序将发出自密钥上次观察到事件以来经过的时间。 然后,它会设置新的计时器,以在 10 秒后发出更新。
要完整运行该示例,需要预先写入一条传感器读数作为流式数据源。 由于定时器使用处理时间,驱动程序会使用 processingTime 触发器,并在停止查询前先等待,以便让定时器触发。
Python
import datetime
import time
import uuid
import pandas as pd
from pyspark.sql.streaming import StatefulProcessor, StatefulProcessorHandle
from pyspark.sql.types import StructType, StructField, StringType, TimestampType
from typing import Iterator
spark.conf.set("spark.sql.streaming.stateStore.providerClass", "org.apache.spark.sql.execution.streaming.state.RocksDBStateStoreProvider")
class DownTimeDetectorStatefulProcessor(StatefulProcessor):
def init(self, handle: StatefulProcessorHandle) -> None:
# Define the schema for the state value (timestamp)
state_schema = StructType([StructField("value", TimestampType(), True)])
self.handle = handle
# Initialize state to store the last seen timestamp for each key
self.last_seen = handle.getValueState("last_seen", state_schema)
def handleExpiredTimer(self, key, timerValues, expiredTimerInfo) -> Iterator[pd.DataFrame]:
latest_from_existing = self.last_seen.get()
# Calculate downtime as the elapsed time between the last observed event and now
downtime_duration = timerValues.getCurrentProcessingTimeInMs() - int(latest_from_existing[0].timestamp() * 1000)
# Register a new timer for 10 seconds in the future
self.handle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + 10000)
# Yield a DataFrame with the key and downtime duration
yield pd.DataFrame(
{
"id": key,
"timeValues": str(downtime_duration),
}
)
def handleInputRows(self, key, rows, timerValues) -> Iterator[pd.DataFrame]:
# Find the row with the maximum timestamp
max_row = max((tuple(pdf.iloc[0]) for pdf in rows), key=lambda row: row[1])
# Get the latest timestamp from the existing state or use epoch start if a timestamp doesn't exist
if self.last_seen.exists():
latest_from_existing = self.last_seen.get()[0]
else:
latest_from_existing = datetime.datetime.fromtimestamp(0)
# If the new data is more recent than the existing state
if latest_from_existing < max_row[1]:
# Delete all existing timers
for timer in self.handle.listTimers():
self.handle.deleteTimer(timer)
# Update the last seen timestamp
self.last_seen.update((max_row[1],))
# Register a new timer for 5 seconds in the future
self.handle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + 5000)
# Get current processing time in milliseconds
timestamp_in_millis = str(timerValues.getCurrentProcessingTimeInMs())
# Yield a DataFrame with the key and current timestamp
yield pd.DataFrame({"id": key, "timeValues": timestamp_in_millis})
def close(self) -> None:
# No cleanup needed
pass
# Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
# Seed a small Delta table with a sensor reading to use as the streaming source
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.sensor_events")
spark.createDataFrame(
[("sensor1", datetime.datetime(2024, 1, 1, 12, 0, 0))],
"id string, timestamp timestamp",
).write.saveAsTable("main.stateful_examples.sensor_events")
df = spark.readStream.table("main.stateful_examples.sensor_events")
# Output schema: the key and a time value (processing time or elapsed downtime)
output_schema = StructType([
StructField("id", StringType(), True),
StructField("timeValues", StringType(), True),
])
# ProcessingTime mode enables the timers that detect downtime
q = (
df.groupBy("id")
.transformWithStateInPandas(
statefulProcessor=DownTimeDetectorStatefulProcessor(),
outputStructType=output_schema,
outputMode="Update",
timeMode="ProcessingTime",
)
.writeStream.format("memory")
.queryName("downtime_output")
.option("checkpointLocation", f"/tmp/checkpoint_{uuid.uuid4()}")
.trigger(processingTime="5 seconds")
.start()
)
# Wait past the timers so they fire, then stop the query
time.sleep(30)
q.stop()
# When a timer fires, it emits the elapsed time in milliseconds since the last observed event
display(spark.sql("SELECT * FROM downtime_output"))
Scala(编程语言)
import java.sql.Timestamp
import org.apache.spark.sql.Encoders
import org.apache.spark.sql.streaming._
import spark.implicits._
import java.util.UUID
// The (String, Timestamp) schema represents an (id, time). We want to do downtime
// detection on every single unique sensor, where each sensor has a sensor ID.
// downtimeThresholdMs is the timer duration in milliseconds.
class DowntimeDetector(downtimeThresholdMs: Long) extends
StatefulProcessor[String, (String, Timestamp), (String, Long)] {
@transient private var _lastSeen: ValueState[Timestamp] = _
private val timestampEncoder = Encoders.TIMESTAMP
override def init(outputMode: OutputMode, timeMode: TimeMode): Unit = {
_lastSeen = getHandle.getValueState[Timestamp]("lastSeen", timestampEncoder, TTLConfig.NONE)
}
// The logic here is as follows: find the largest timestamp seen so far. Set a timer for
// the duration later.
override def handleInputRows(
key: String,
inputRows: Iterator[(String, Timestamp)],
timerValues: TimerValues): Iterator[(String, Long)] = {
val latestRecordFromNewRows = inputRows.maxBy(_._2.getTime)
// Use getOrElse to initiate state variable if it doesn't exist
val latestTimestampFromExistingRows = Option(_lastSeen.get()).getOrElse(new Timestamp(0))
val latestTimestampFromNewRows = latestRecordFromNewRows._2
if (latestTimestampFromNewRows.after(latestTimestampFromExistingRows)) {
// Cancel the one existing timer, since we have a new latest timestamp.
// We call "listTimers()" because we don't know ahead of time what
// the timestamp of the existing timer will be.
getHandle.listTimers().foreach(timer => getHandle.deleteTimer(timer))
_lastSeen.update(latestTimestampFromNewRows)
// Use timerValues to schedule a timer using processing time.
getHandle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + downtimeThresholdMs)
} else {
// No new latest timestamp, so there is no need to update the state or set a timer.
}
Iterator.empty
}
override def handleExpiredTimer(
key: String,
timerValues: TimerValues,
expiredTimerInfo: ExpiredTimerInfo): Iterator[(String, Long)] = {
val latestTimestamp = _lastSeen.get()
// Downtime is the elapsed time in milliseconds between the last observed event and now
val downtimeDurationMs =
timerValues.getCurrentProcessingTimeInMs() - latestTimestamp.getTime
// Register another timer that will fire in 10 seconds.
// Timers can be registered anywhere but init()
getHandle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + 10000)
Iterator((key, downtimeDurationMs))
}
}
// Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
// Seed a small Delta table with a sensor reading to use as the streaming source
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.sensor_events_scala")
Seq(
("sensor1", Timestamp.valueOf("2024-01-01 12:00:00"))
).toDF("id", "timestamp").write.saveAsTable("main.stateful_examples.sensor_events_scala")
// ProcessingTime mode enables the timers that detect downtime
val q = spark.readStream
.table("main.stateful_examples.sensor_events_scala")
.as[(String, Timestamp)]
.groupByKey(_._1)
.transformWithState(
new DowntimeDetector(5000L),
TimeMode.ProcessingTime(),
OutputMode.Update()
)
.writeStream
.format("memory")
.queryName("downtime_output_scala")
.option("checkpointLocation", s"/tmp/checkpoint_${UUID.randomUUID()}")
.trigger(Trigger.ProcessingTime("5 seconds"))
.start()
// Wait past the timers so they fire, then stop the query
Thread.sleep(30000)
q.stop()
// When a timer fires, it emits the elapsed time in milliseconds since the last observed event
spark.sql("SELECT * FROM downtime_output_scala").show(false)
迁移现有状态信息
以下示例演示如何实现接受初始状态的有状态应用程序。 可以将初始状态处理添加到任何有状态应用程序,但初始状态只能在首次初始化应用程序时设置。
此示例使用 statestore 读取器从检查点路径加载现有状态信息。 此模式的示例用例是从旧有状态应用程序迁移到 transformWithState的。
Python
# Import the necessary libraries
import pandas as pd
from pyspark.sql.streaming import StatefulProcessor, StatefulProcessorHandle
from pyspark.sql.types import StructType, StructField, LongType, StringType, IntegerType
from typing import Iterator
# Set RocksDB as the state store provider for better performance
spark.conf.set("spark.sql.streaming.stateStore.providerClass", "org.apache.spark.sql.execution.streaming.state.RocksDBStateStoreProvider")
"""
Input schema is as below
input_schema = StructType(
[StructField("id", StringType(), True)],
[StructField("value", StringType(), True)]
)
"""
# Define the output schema for the streaming query
output_schema = StructType([
StructField("id", StringType(), True),
StructField("accumulated", StringType(), True)
])
class AccumulatedCounterStatefulProcessorWithInitialState(StatefulProcessor):
def init(self, handle: StatefulProcessorHandle) -> None:
# Define the schema for the state value (integer)
state_schema = StructType([StructField("value", IntegerType(), True)])
# Initialize state to store the accumulated counter for each id
self.counter_state = handle.getValueState("counter_state", state_schema)
self.handle = handle
def handleInputRows(self, key, rows, timerValues) -> Iterator[pd.DataFrame]:
# Check if state exists for the current key
exists = self.counter_state.exists()
if exists:
value_row = self.counter_state.get()
existing_value = value_row[0]
else:
existing_value = 0
accumulated_value = existing_value
# Process input rows and accumulate values
for pdf in rows:
value = pdf["value"].astype(int).sum()
accumulated_value += value
# Update the state with the new accumulated value
self.counter_state.update((accumulated_value,))
# Yield a DataFrame with the key and accumulated value
yield pd.DataFrame({"id": key, "accumulated": str(accumulated_value)})
def handleInitialState(self, key, initialState, timerValues) -> None:
# Initialize the state with the provided initial value
init_val = initialState.at[0, "initVal"]
self.counter_state.update((init_val,))
def close(self) -> None:
# No cleanup needed
pass
# Load initial state from a checkpoint directory
initial_state = spark.read.format("statestore")
.option("path", "$checkpointsDir")
.load()
# Apply the stateful transformation to the input DataFrame
df.groupBy("id")
.transformWithStateInPandas(
statefulProcessor=AccumulatedCounterStatefulProcessorWithInitialState(),
outputStructType=output_schema,
outputMode="Update",
timeMode="None",
initialState=initial_state,
)
.writeStream... # Continue with stream writing configuration
Scala(编程语言)
// Import the necessary libraries
import org.apache.spark.sql.streaming._
import org.apache.spark.sql.{Dataset, Encoder, Encoders, DataFrame}
import org.apache.spark.sql.types._
// Define a stateful processor that can handle the initial state
class InitialStateStatefulProcessor extends StatefulProcessorWithInitialState[String, (String, String, String), (String, String), (String, Int)] {
// Transient value state to store the accumulated value
@transient protected var valueState: ValueState[Int] = _
private val intEncoder = Encoders.scalaInt
// Initialize the state store
override def init(
outputMode: OutputMode,
timeMode: TimeMode): Unit = {
// Create a value state named "valueState" using Int encoder
// TTLConfig.NONE means the state has no automatic expiration
valueState = getHandle.getValueState[Int]("valueState",
intEncoder, TTLConfig.NONE)
}
// Process input rows and update state
override def handleInputRows(
key: String,
inputRows: Iterator[(String, String, String)],
timerValues: TimerValues): Iterator[(String, String)] = {
var existingValue = 0
// Retrieve existing value from state if it exists
if (valueState.exists()) {
existingValue += valueState.get()
}
var accumulatedValue = existingValue
// Accumulate values from input rows
for (row <- inputRows) {
accumulatedValue += row._2.toInt
}
// Update the state with the new accumulated value
valueState.update(accumulatedValue)
// Return the key and accumulated value as a string
Iterator((key, accumulatedValue.toString))
}
// Handle initial state when provided
override def handleInitialState(
key: String, initialState: (String, Int), timerValues: TimerValues): Unit = {
// Update the state with the initial value
valueState.update(initialState._2)
}
}
将 Delta 表迁移到用于初始化的状态存储
以下笔记本包含一个在 Python 或 Scala 中使用 Delta 表 transformWithState 初始化状态存储值的示例。
从 Delta Python 初始化状态
从 Delta Scala 初始化状态
会话跟踪
以下笔记本包含在 Python 或 Scala 中使用 transformWithState 的会话跟踪示例。
会话跟踪 Python
会话跟踪 Scala
使用 transformWithState 的自定义流-流联接
以下代码演示了使用 transformWithState 跨多个流的自定义流-流联接。 出于以下原因,可以使用此方法而不是内置联接运算符:
- 需要使用不支持流-流联接的更新输出模式。 这对于较低的延迟应用程序尤其有用。
- 需要继续对延迟到达行执行联接(水印过期之后)。
- 需要执行多对多流-流联接。
这个示例使你能够完全控制状态过期逻辑,并支持动态延长保留期,以便在水印之后仍能处理乱序事件。
在以下示例中,配置、偏好和活动事件都通过同一个流到达,并且每个事件都带有 record_type 标签。 处理器将每种记录类型缓存在状态中,并由处理时间定时器在活动事件到达后不久输出增强后的连接结果。 使用TTL后,配置文件和偏好状态在一小时未激活后失效,每个活动一旦加入后即被清除。
注释
这个示例为每个用户保留一条活动记录,并在 join 产生输出后将其清除。 为了保持专注,它不会处理同一个用户在计时器触发前收到的多个活动事件:一个较晚的活动替换了之前的,每个计时器读取的是最新的缓冲活动,而不是调度后的那个。 为了保留每个活动,请将活动缓冲到以事件时间为键的列表状态或映射状态中。
Python
# Import the necessary libraries
import pandas as pd
import time
import uuid
from datetime import datetime
from pyspark.sql.streaming import StatefulProcessor, StatefulProcessorHandle
from pyspark.sql.types import StructType, StructField, StringType, TimestampType
from typing import Iterator
spark.conf.set("spark.sql.streaming.stateStore.providerClass", "org.apache.spark.sql.execution.streaming.state.RocksDBStateStoreProvider")
# Define output schema for the joined data
output_schema = StructType([
StructField("user_id", StringType(), True),
StructField("event_type", StringType(), True),
StructField("timestamp", TimestampType(), True),
StructField("profile_name", StringType(), True),
StructField("email", StringType(), True),
StructField("preferred_category", StringType(), True)
])
class CustomStreamJoinProcessor(StatefulProcessor):
# Buffer each user's profile, preference, and activity records in state.
def init(self, handle: StatefulProcessorHandle) -> None:
self.handle = handle
profile_schema = StructType([
StructField("name", StringType(), True),
StructField("email", StringType(), True)
])
preferences_schema = StructType([
StructField("preferred_category", StringType(), True)
])
activity_schema = StructType([
StructField("event_type", StringType(), True),
StructField("timestamp", TimestampType(), True)
])
# One value state per record type. The grouping key is user_id, so each
# state holds the latest record of that type for the user.
# Profile and preference state expire after an hour of inactivity via TTL
self.profile_state = handle.getValueState("userProfile", profile_schema, ttlDurationMs=3600000)
self.preferences_state = handle.getValueState("userPreferences", preferences_schema, ttlDurationMs=3600000)
self.activity_state = handle.getValueState("userActivity", activity_schema)
# Route each incoming record by its type and buffer it in state. When an
# activity event arrives, set a timer to emit the enriched join after a delay.
def handleInputRows(self, key, rows: Iterator[pd.DataFrame], timerValues) -> Iterator[pd.DataFrame]:
for pdf in rows:
for _, row in pdf.iterrows():
record_type = row["record_type"]
if record_type == "activity":
self.activity_state.update((row["event_type"], row["timestamp"]))
# Set a timer to process this event after a 10-second delay
self.handle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + 10000)
elif record_type == "profile":
self.profile_state.update((row["name"], row["email"]))
elif record_type == "preference":
self.preferences_state.update((row["preferred_category"],))
# No immediate output; the enriched row is emitted when the timer expires
return iter([])
# Perform the lookup after the delay, handling out-of-order and late-arriving records.
def handleExpiredTimer(self, key, timerValues, expiredTimerInfo) -> Iterator[pd.DataFrame]:
if not self.activity_state.exists():
return iter([])
activity = self.activity_state.get()
profile = self.profile_state.get() if self.profile_state.exists() else None
preferences = self.preferences_state.get() if self.preferences_state.exists() else None
# Combine data from the different states into a single output row
output_row = {
"user_id": key[0],
"event_type": activity[0],
"timestamp": activity[1],
"profile_name": profile[0] if profile else None,
"email": profile[1] if profile else None,
"preferred_category": preferences[0] if preferences else None
}
# The activity has been consumed by this join, so clear it from state
self.activity_state.clear()
return iter([pd.DataFrame([output_row])])
def close(self) -> None:
pass
# Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
# Seed a small Delta table with profile, preference, and activity records for one user
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.user_events")
input_schema = StructType([
StructField("user_id", StringType()),
StructField("record_type", StringType()),
StructField("event_type", StringType()),
StructField("timestamp", TimestampType()),
StructField("name", StringType()),
StructField("email", StringType()),
StructField("preferred_category", StringType())
])
spark.createDataFrame(
[
("u1", "profile", None, None, "Alice", "alice@example.com", None),
("u1", "preference", None, None, None, None, "electronics"),
("u1", "activity", "purchase", datetime(2024, 1, 1, 12, 0, 0), None, None, None),
],
input_schema,
).write.saveAsTable("main.stateful_examples.user_events")
df = spark.readStream.table("main.stateful_examples.user_events")
# Apply transformWithState. ProcessingTime mode enables the timer that fires the join.
q = (
df.groupBy("user_id")
.transformWithStateInPandas(
statefulProcessor=CustomStreamJoinProcessor(),
outputStructType=output_schema,
outputMode="Append",
timeMode="ProcessingTime",
)
.writeStream.format("memory")
.queryName("enriched_events")
.option("checkpointLocation", f"/tmp/checkpoint_{uuid.uuid4()}")
.trigger(processingTime="5 seconds")
.start()
)
# Wait past the 10-second timer so it fires, then stop the query
time.sleep(30)
q.stop()
# The enriched row joins the activity with the buffered profile and preference
display(spark.sql("SELECT * FROM enriched_events"))
Scala(编程语言)
// Import the necessary libraries
import org.apache.spark.sql.streaming._
import org.apache.spark.sql.Encoders
import spark.implicits._
import java.sql.Timestamp
import java.util.UUID
import java.time.Duration
// Unified input record: every event arrives on one stream, tagged by record_type
case class UserRecord(
user_id: String,
record_type: String,
event_type: Option[String],
timestamp: Option[Timestamp],
name: Option[String],
email: Option[String],
preferred_category: Option[String]
)
case class UserActivity(event_type: String, timestamp: Timestamp)
case class UserProfile(name: String, email: String)
case class UserPreferences(preferred_category: String)
// Enriched user event combining activity with profile and preference data
case class EnrichedUserEvent(
user_id: String,
event_type: String,
timestamp: Timestamp,
profile_name: Option[String],
email: Option[String],
preferred_category: Option[String]
)
// Custom stateful processor for the stream-stream join
class CustomStreamJoinProcessor extends StatefulProcessor[String, UserRecord, EnrichedUserEvent] {
// One value state per record type. The grouping key is user_id, so each state
// holds the latest record of that type for the user.
@transient private var _profileState: ValueState[UserProfile] = _
@transient private var _preferencesState: ValueState[UserPreferences] = _
@transient private var _activityState: ValueState[UserActivity] = _
override def init(outputMode: OutputMode, timeMode: TimeMode): Unit = {
// Profile and preference state expire after an hour of inactivity via TTL
_profileState = getHandle.getValueState[UserProfile]("profileState", Encoders.product[UserProfile], TTLConfig(Duration.ofHours(1)))
_preferencesState = getHandle.getValueState[UserPreferences]("preferencesState", Encoders.product[UserPreferences], TTLConfig(Duration.ofHours(1)))
_activityState = getHandle.getValueState[UserActivity]("activityState", Encoders.product[UserActivity], TTLConfig.NONE)
}
// Route each incoming record by its type and buffer it in state. When an
// activity event arrives, set a timer to emit the enriched join after a delay.
override def handleInputRows(
key: String,
inputRows: Iterator[UserRecord],
timerValues: TimerValues): Iterator[EnrichedUserEvent] = {
inputRows.foreach { rec =>
rec.record_type match {
case "activity" =>
_activityState.update(UserActivity(rec.event_type.getOrElse(""), rec.timestamp.orNull))
getHandle.registerTimer(timerValues.getCurrentProcessingTimeInMs() + 10000)
case "profile" =>
_profileState.update(UserProfile(rec.name.getOrElse(""), rec.email.getOrElse("")))
case "preference" =>
_preferencesState.update(UserPreferences(rec.preferred_category.getOrElse("")))
case _ =>
}
}
Iterator.empty
}
// When the timer expires, join the buffered activity with the latest profile and preference
override def handleExpiredTimer(
key: String,
timerValues: TimerValues,
expiredTimerInfo: ExpiredTimerInfo): Iterator[EnrichedUserEvent] = {
if (!_activityState.exists()) {
Iterator.empty
} else {
val activity = _activityState.get()
val profile = if (_profileState.exists()) Some(_profileState.get()) else None
val preferences = if (_preferencesState.exists()) Some(_preferencesState.get()) else None
// The activity has been consumed by this join, so clear it from state
_activityState.clear()
Iterator.single(EnrichedUserEvent(
user_id = key,
event_type = activity.event_type,
timestamp = activity.timestamp,
profile_name = profile.map(_.name),
email = profile.map(_.email),
preferred_category = preferences.map(_.preferred_category)
))
}
}
}
// Create a dedicated schema for the example tables
spark.sql("CREATE SCHEMA IF NOT EXISTS main.stateful_examples")
// Seed a small Delta table with profile, preference, and activity records for one user
spark.sql("DROP TABLE IF EXISTS main.stateful_examples.user_events_scala")
Seq(
UserRecord("u1", "profile", None, None, Some("Alice"), Some("alice@example.com"), None),
UserRecord("u1", "preference", None, None, None, None, Some("electronics")),
UserRecord("u1", "activity", Some("purchase"), Some(Timestamp.valueOf("2024-01-01 12:00:00")), None, None, None)
).toDF().write.saveAsTable("main.stateful_examples.user_events_scala")
// Apply the custom stateful processor. ProcessingTime mode enables the join timer.
val enrichedStream = spark.readStream
.table("main.stateful_examples.user_events_scala")
.as[UserRecord]
.groupByKey(_.user_id)
.transformWithState(
new CustomStreamJoinProcessor(),
TimeMode.ProcessingTime(),
OutputMode.Append()
)
val q = enrichedStream.writeStream
.format("memory")
.queryName("enriched_events_scala")
.option("checkpointLocation", s"/tmp/checkpoint_${UUID.randomUUID()}")
.trigger(Trigger.ProcessingTime("5 seconds"))
.start()
// Wait past the 10-second timer so it fires, then stop the query
Thread.sleep(30000)
q.stop()
// The enriched row joins the activity with the buffered profile and preference
spark.sql("SELECT * FROM enriched_events_scala").show(false)
Top-K 计算
以下示例使用具有优先级队列的 ListState,以近乎实时地维护和更新每个组键的数据流中的前 K 个元素。