Raspberry Pi、Kinesis Video Streams、AWS Rekognitionによるリアルタイム顔検出

Raspberry Pi、Kinesis Video Streams、AWS Rekognitionによるリアルタイム顔検出

Raspberry Pi上でKinesis Video StreamsとAmazon Rekognition Videoを使い、リアルタイム顔検出を構築する。

Takahiro Iwasa
10 min read

Raspberry PiにUSBカメラを接続し、Kinesis Video StreamsRekognition Videoを組み合わせることで、リアルタイムに顔を検出する顔検出システムを構築できます。

必要要件

ハードウェア要件

  • Raspberry Pi 4B(RAM 4GB)
  • USBカメラ

ソフトウェア要件

AWSリソースの構築

AWS SAMテンプレート

template.yaml
AWSTemplateFormatVersion: 2010-09-09
Transform: AWS::Serverless-2016-10-31
Description: face-detector-using-kinesis-video-streams
Resources:
Function:
Type: AWS::Serverless::Function
Properties:
FunctionName: face-detector-function
CodeUri: src/
Handler: app.lambda_handler
Runtime: python3.11
Architectures:
- arm64
Timeout: 3
MemorySize: 128
Role: !GetAtt FunctionIAMRole.Arn
Events:
KinesisEvent:
Type: Kinesis
Properties:
Stream: !GetAtt KinesisStream.Arn
MaximumBatchingWindowInSeconds: 10
MaximumRetryAttempts: 3
StartingPosition: LATEST
FunctionIAMRole:
Type: AWS::IAM::Role
Properties:
RoleName: face-detector-function-role
AssumeRolePolicyDocument:
Version: 2012-10-17
Statement:
- Effect: Allow
Principal:
Service: lambda.amazonaws.com
Action: sts:AssumeRole
ManagedPolicyArns:
- arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole
- arn:aws:iam::aws:policy/service-role/AWSLambdaKinesisExecutionRole
Policies:
- PolicyName: policy
PolicyDocument:
Version: 2012-10-17
Statement:
- Effect: Allow
Action:
- kinesisvideo:GetHLSStreamingSessionURL
- kinesisvideo:GetDataEndpoint
Resource: !GetAtt KinesisVideoStream.Arn
KinesisVideoStream:
Type: AWS::KinesisVideo::Stream
Properties:
Name: face-detector-kinesis-video-stream
DataRetentionInHours: 24
RekognitionCollection:
Type: AWS::Rekognition::Collection
Properties:
CollectionId: FaceCollection
RekognitionStreamProcessor:
Type: AWS::Rekognition::StreamProcessor
Properties:
Name: face-detector-rekognition-stream-processor
KinesisVideoStream:
Arn: !GetAtt KinesisVideoStream.Arn
KinesisDataStream:
Arn: !GetAtt KinesisStream.Arn
RoleArn: !GetAtt RekognitionStreamProcessorIAMRole.Arn
FaceSearchSettings:
CollectionId: !Ref RekognitionCollection
FaceMatchThreshold: 80
DataSharingPreference:
OptIn: false
KinesisStream:
Type: AWS::Kinesis::Stream
Properties:
Name: face-detector-kinesis-stream
StreamModeDetails:
StreamMode: ON_DEMAND
RekognitionStreamProcessorIAMRole:
Type: AWS::IAM::Role
Properties:
RoleName: face-detector-rekognition-stream-processor-role
AssumeRolePolicyDocument:
Version: 2012-10-17
Statement:
- Effect: Allow
Principal:
Service: rekognition.amazonaws.com
Action: sts:AssumeRole
ManagedPolicyArns:
- arn:aws:iam::aws:policy/service-role/AmazonRekognitionServiceRole
Policies:
- PolicyName: policy
PolicyDocument:
Version: 2012-10-17
Statement:
- Effect: Allow
Action:
- kinesis:PutRecord
- kinesis:PutRecords
Resource:
- !GetAtt KinesisStream.Arn

Lambda関数

src/app.py
import base64
import json
import logging
from datetime import datetime, timedelta, timezone
from functools import cache
import boto3
JST = timezone(timedelta(hours=9))
kvs_client = boto3.client('kinesisvideo')
logger = logging.getLogger(__name__)
logger.setLevel(logging.INFO)
def lambda_handler(event: dict, context: dict) -> dict:
for record in event['Records']:
base64_data = record['kinesis']['data']
stream_processor_event = json.loads(base64.b64decode(base64_data).decode())
# Refer to https://docs.aws.amazon.com/rekognition/latest/dg/streaming-video-kinesis-output.html for details on the structure.
if not stream_processor_event['FaceSearchResponse']:
continue
logger.info(stream_processor_event)
url = get_hls_streaming_session_url(stream_processor_event)
logger.info(url)
return {
'statusCode': 200,
}
@cache
def get_kvs_am_client(api_name: str, stream_arn: str):
# Retrieves the data endpoint for the stream.
# See https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/kinesisvideo/client/get_data_endpoint.html
endpoint = kvs_client.get_data_endpoint(
APIName=api_name.upper(),
StreamARN=stream_arn
)['DataEndpoint']
return boto3.client('kinesis-video-archived-media', endpoint_url=endpoint)
def get_hls_streaming_session_url(stream_processor_event: dict) -> str:
# Generates an HLS streaming URL for the video stream.
# See https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/kinesis-video-archived-media/client/get_hls_streaming_session_url.html
kinesis_video = stream_processor_event['InputInformation']['KinesisVideo']
stream_arn = kinesis_video['StreamArn']
kvs_am_client = get_kvs_am_client('get_hls_streaming_session_url', stream_arn)
start_timestamp = datetime.fromtimestamp(kinesis_video['ServerTimestamp'], JST)
end_timestamp = datetime.fromtimestamp(kinesis_video['ServerTimestamp'], JST) + timedelta(minutes=1)
return kvs_am_client.get_hls_streaming_session_url(
StreamARN=stream_arn,
PlaybackMode='ON_DEMAND',
HLSFragmentSelector={
'FragmentSelectorType': 'SERVER_TIMESTAMP',
'TimestampRange': {
'StartTimestamp': start_timestamp,
'EndTimestamp': end_timestamp,
},
},
ContainerFormat='FRAGMENTED_MP4',
Expires=300,
)['HLSStreamingSessionURL']

スタックのデプロイ

SAMアプリケーションをビルドしてデプロイします。

Terminal window
sam build
sam deploy

顔のインデックス登録

USBカメラで顔を検出するには、事前にRekognitionの顔コレクションに顔をインデックス(登録)しておく必要があります。この目的にはIndexFaces APIを使用します。

以下を実際の値に置き換えてください。

  • <YOUR_BUCKET>
  • <YOUR_OBJECT>
  • <PERSON_ID>
Terminal window
aws rekognition index-faces \
--image '{"S3Object": {"Bucket": "<YOUR_BUCKET>", "Name": "<YOUR_OBJECT>"}}' \
--collection-id FaceCollection \
--external-image-id <PERSON_ID>

Rekognitionは顔コレクションに実際の画像を保存しません。代わりに、顔の特徴をメタデータとして抽出し保存します。

https://docs.aws.amazon.com/rekognition/latest/dg/add-faces-to-collection-procedure.html

For each face detected, Amazon Rekognition extracts facial features and stores the feature information in a database. In addition, the command stores metadata for each face that’s detected in the specified face collection. Amazon Rekognition doesn’t store the actual image bytes.

ビデオプロデューサーのセットアップ

この例では、Ubuntu 23.10を実行するRaspberry Pi 4B(RAM 4GB)をビデオプロデューサーとして使用します。

Raspberry Pi and USB Camera Setup

GStreamerプラグインのビルド

AWSはAmazon Kinesis Video Streams CPP Producer, GStreamer Plugin and JNIを提供しています。このSDKにより、Raspberry PiからKinesis Video Streamsへの動画ストリーミングが可能になります。

ℹ️ Note

AWSはGStreamerプラグイン用のDockerイメージを提供していますが、アーキテクチャの制約によりRaspberry Piでは動作しない場合があります。

以下のコマンドを実行します。システムのスペックによっては、ビルドに20分以上かかることがあります。

Terminal window
sudo apt update
sudo apt upgrade
sudo apt install \
make \
cmake \
build-essential \
m4 \
autoconf \
default-jdk
sudo apt install \
libssl-dev \
libcurl4-openssl-dev \
liblog4cplus-dev \
libgstreamer1.0-dev \
libgstreamer-plugins-base1.0-dev \
gstreamer1.0-plugins-base-apps \
gstreamer1.0-plugins-bad \
gstreamer1.0-plugins-good \
gstreamer1.0-plugins-ugly \
gstreamer1.0-tools
git clone https://github.com/awslabs/amazon-kinesis-video-streams-producer-sdk-cpp.git
mkdir -p amazon-kinesis-video-streams-producer-sdk-cpp/build
cd amazon-kinesis-video-streams-producer-sdk-cpp/build
sudo cmake .. -DBUILD_GSTREAMER_PLUGIN=ON -DBUILD_JNI=TRUE
sudo make

ビルドが完了したら、以下のコマンドで結果を確認します。

Terminal window
cd ~/amazon-kinesis-video-streams-producer-sdk-cpp
export GST_PLUGIN_PATH=`pwd`/build
export LD_LIBRARY_PATH=`pwd`/open-source/local/lib
gst-inspect-1.0 kvssink

以下のような出力が表示されるはずです。

Factory Details:
Rank primary + 10 (266)
Long-name KVS Sink
Klass Sink/Video/Network
Description GStreamer AWS KVS plugin
Author AWS KVS <[email protected]>
...

毎回環境変数をリセットしなくて済むように、~/.profileに以下のexport文を追加します。

Terminal window
echo "" >> ~/.profile
echo "# GStreamer" >> ~/.profile
echo "export GST_PLUGIN_PATH=$GST_PLUGIN_PATH" >> ~/.profile
echo "export LD_LIBRARY_PATH=$LD_LIBRARY_PATH" >> ~/.profile

GStreamerの実行

プラグインをビルドしたら、USBカメラをRaspberry Piに接続し、以下のコマンドを実行して動画データをKinesis Video Streamsにストリーミングします。

以下を実際の値に置き換えてください。

  • <KINESIS_VIDEO_STREAM_NAME>
  • <YOUR_ACCESS_KEY>
  • <YOUR_SECRET_KEY>
  • <YOUR_AWS_REGION>
🔥 Caution

動画品質を上げる(解像度やフレームレートを上げるなど)と、AWSのコストが増加する可能性があります。

Terminal window
gst-launch-1.0 -v v4l2src device=/dev/video0 \
! videoconvert \
! video/x-raw,format=I420,width=320,height=240,framerate=5/1 \
! x264enc bframes=0 key-int-max=45 bitrate=500 tune=zerolatency \
! video/x-h264,stream-format=avc,alignment=au \
! kvssink stream-name=<KINESIS_VIDEO_STREAM_NAME> storage-size=128 access-key="<YOUR_ACCESS_KEY>" secret-key="<YOUR_SECRET_KEY>" aws-region="<YOUR_AWS_REGION>"

Kinesis Video Streams管理コンソールに移動すると、ライブストリームを確認できます。

Kinesis Video Streams Management Console

テスト

Rekognition Videoストリームプロセッサの起動

Rekognition Videoストリームプロセッサを起動します。このサービスはKinesis Video Streamを購読し、顔コレクションを使って顔を検出し、その結果をKinesis Data Streamにストリーミングします。

ストリームプロセッサを起動します。

Terminal window
aws rekognition start-stream-processor \
--name face-detector-rekognition-stream-processor

ストリームプロセッサが実行中であることを確認するため、ステータスを確認します。

Terminal window
aws rekognition describe-stream-processor \
--name face-detector-rekognition-stream-processor | grep "Status"

"Status": "RUNNING"と表示されれば正常です。

顔の検出

USBカメラが動画をキャプチャすると、Rekognition Videoストリームプロセッサが動画ストリームを解析し、顔コレクションに基づいて顔を検出します。

結果を確認するには、以下のコマンドでLambda関数のログを表示します。

Terminal window
sam logs -n Function \
--stack-name face-detector-using-kinesis-video-streams \
--tail

ログレコードには、以下の例のようにストリームプロセッサイベントの詳細情報が含まれます。

{
"InputInformation": {
"KinesisVideo": {
"StreamArn": "arn:aws:kinesisvideo:<AWS_REGION>:<AWS_ACCOUNT_ID>:stream/face-detector-kinesis-video-stream/xxxxxxxxxxxxx",
"FragmentNumber": "91343852333181501717324262640137742175000164731",
"ServerTimestamp": 1702208586.022,
"ProducerTimestamp": 1702208585.699,
"FrameOffsetInSeconds": 0.0,
}
},
"StreamProcessorInformation": {"Status": "RUNNING"},
"FaceSearchResponse": [
{
"DetectedFace": {
"BoundingBox": {
"Height": 0.4744676,
"Width": 0.29107505,
"Left": 0.33036956,
"Top": 0.19599175,
},
"Confidence": 99.99677,
"Landmarks": [
{"X": 0.41322955, "Y": 0.33761832, "Type": "eyeLeft"},
{"X": 0.54405355, "Y": 0.34024307, "Type": "eyeRight"},
{"X": 0.424819, "Y": 0.5417343, "Type": "mouthLeft"},
{"X": 0.5342691, "Y": 0.54362005, "Type": "mouthRight"},
{"X": 0.48934412, "Y": 0.43806323, "Type": "nose"},
],
"Pose": {"Pitch": 5.547308, "Roll": 0.85795176, "Yaw": 4.76913},
"Quality": {"Brightness": 57.938313, "Sharpness": 46.0298},
},
"MatchedFaces": [
{
"Similarity": 99.986176,
"Face": {
"BoundingBox": {
"Height": 0.417963,
"Width": 0.406223,
"Left": 0.28826,
"Top": 0.242463,
},
"FaceId": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"Confidence": 99.996605,
"ImageId": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
"ExternalImageId": "iwasa",
},
}
],
}
],
}

動画再生用のHLS URL

ログには、オンデマンド動画再生用に生成されたHLS URLも含まれます。例えば以下のようなものです。

https://x-xxxxxxxx.kinesisvideo.<AWS_REGION>.amazonaws.com/hls/v1/getHLSMasterPlaylist.m3u8?SessionToken=xxxxxxxxxx

SafariやEdgeなど、対応しているブラウザでHLS URLを開いてください。

ℹ️ Note

ChromeはHLS再生をネイティブでサポートしていません。Native HLS Playbackのようなサードパーティ拡張機能を使用できます。

HLS Playback Example

クリーンアップ

作業が終わったら、ストリームプロセッサを停止しスタックを削除します。

Terminal window
aws rekognition stop-stream-processor \
--name face-detector-rekognition-stream-processor
sam delete

まとめ

Raspberry PiのUSBカメラ映像をKinesis Video StreamsとRekognition Videoに流したところ、事前にインデックス登録した顔コレクションに対するリアルタイムの顔マッチが得られ、Lambda関数が各マッチを再生可能なHLS URLに変換しました。このパイプラインが独自の動画処理コードなしに機能するのは、Rekognition Videoのストリームプロセッサが、Kinesis Video Streamに対して直接顔の検出とマッチングを行っているためであり、Lambda関数には上記のとおり比較的シンプルな役割しか残っていません。とはいえ、kvssinkに渡すエンコード設定(解像度、フレームレート、ビットレート)はAWSのコストに直結するため、まずは本記事で使用した低い値から始め、検出精度が実際に必要とする場合にのみ引き上げる方が、最初から高品質をデフォルトにするより望ましいでしょう。

About the author

Takahiro Iwasa

Takahiro Iwasa

Software Developer

This blog shares technical notes from hands-on projects—architecture, implementation, and AWS service integrations.