Raspberry Pi、Kinesis Video Streams、AWS Rekognitionによるリアルタイム顔検出
Raspberry Pi上でKinesis Video StreamsとAmazon Rekognition Videoを使い、リアルタイム顔検出を構築する。
Raspberry PiにUSBカメラを接続し、Kinesis Video StreamsとRekognition Videoを組み合わせることで、リアルタイムに顔を検出する顔検出システムを構築できます。
必要要件
ハードウェア要件
- Raspberry Pi 4B(RAM 4GB)
- Ubuntu 23.10を実行(Raspberry Pi Imagerでインストール)
- USBカメラ
ソフトウェア要件
- GStreamer
- Amazon Kinesis Video Streams CPP Producer, GStreamer Plugin and JNI
- AWS SAM CLI
- Python 3.11
AWSリソースの構築
AWS SAMテンプレート
AWSTemplateFormatVersion: 2010-09-09Transform: AWS::Serverless-2016-10-31Description: face-detector-using-kinesis-video-streams
Resources: Function: Type: AWS::Serverless::Function Properties: FunctionName: face-detector-function CodeUri: src/ Handler: app.lambda_handler Runtime: python3.11 Architectures: - arm64 Timeout: 3 MemorySize: 128 Role: !GetAtt FunctionIAMRole.Arn Events: KinesisEvent: Type: Kinesis Properties: Stream: !GetAtt KinesisStream.Arn MaximumBatchingWindowInSeconds: 10 MaximumRetryAttempts: 3 StartingPosition: LATEST
FunctionIAMRole: Type: AWS::IAM::Role Properties: RoleName: face-detector-function-role AssumeRolePolicyDocument: Version: 2012-10-17 Statement: - Effect: Allow Principal: Service: lambda.amazonaws.com Action: sts:AssumeRole ManagedPolicyArns: - arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole - arn:aws:iam::aws:policy/service-role/AWSLambdaKinesisExecutionRole Policies: - PolicyName: policy PolicyDocument: Version: 2012-10-17 Statement: - Effect: Allow Action: - kinesisvideo:GetHLSStreamingSessionURL - kinesisvideo:GetDataEndpoint Resource: !GetAtt KinesisVideoStream.Arn
KinesisVideoStream: Type: AWS::KinesisVideo::Stream Properties: Name: face-detector-kinesis-video-stream DataRetentionInHours: 24
RekognitionCollection: Type: AWS::Rekognition::Collection Properties: CollectionId: FaceCollection
RekognitionStreamProcessor: Type: AWS::Rekognition::StreamProcessor Properties: Name: face-detector-rekognition-stream-processor KinesisVideoStream: Arn: !GetAtt KinesisVideoStream.Arn KinesisDataStream: Arn: !GetAtt KinesisStream.Arn RoleArn: !GetAtt RekognitionStreamProcessorIAMRole.Arn FaceSearchSettings: CollectionId: !Ref RekognitionCollection FaceMatchThreshold: 80 DataSharingPreference: OptIn: false
KinesisStream: Type: AWS::Kinesis::Stream Properties: Name: face-detector-kinesis-stream StreamModeDetails: StreamMode: ON_DEMAND
RekognitionStreamProcessorIAMRole: Type: AWS::IAM::Role Properties: RoleName: face-detector-rekognition-stream-processor-role AssumeRolePolicyDocument: Version: 2012-10-17 Statement: - Effect: Allow Principal: Service: rekognition.amazonaws.com Action: sts:AssumeRole ManagedPolicyArns: - arn:aws:iam::aws:policy/service-role/AmazonRekognitionServiceRole Policies: - PolicyName: policy PolicyDocument: Version: 2012-10-17 Statement: - Effect: Allow Action: - kinesis:PutRecord - kinesis:PutRecords Resource: - !GetAtt KinesisStream.ArnLambda関数
- Rekognition Videoのストリームプロセッサは、検出した顔のデータをKinesis Data Streamにストリーミングします。データはBase64文字列としてエンコードされています(17行目)。
- Lambda関数はHLS URLを生成します(54〜66行目)。
import base64import jsonimport loggingfrom datetime import datetime, timedelta, timezonefrom functools import cache
import boto3
JST = timezone(timedelta(hours=9))kvs_client = boto3.client('kinesisvideo')logger = logging.getLogger(__name__)logger.setLevel(logging.INFO)
def lambda_handler(event: dict, context: dict) -> dict: for record in event['Records']: base64_data = record['kinesis']['data'] stream_processor_event = json.loads(base64.b64decode(base64_data).decode()) # Refer to https://docs.aws.amazon.com/rekognition/latest/dg/streaming-video-kinesis-output.html for details on the structure.
if not stream_processor_event['FaceSearchResponse']: continue
logger.info(stream_processor_event) url = get_hls_streaming_session_url(stream_processor_event) logger.info(url)
return { 'statusCode': 200, }
@cachedef get_kvs_am_client(api_name: str, stream_arn: str): # Retrieves the data endpoint for the stream. # See https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/kinesisvideo/client/get_data_endpoint.html endpoint = kvs_client.get_data_endpoint( APIName=api_name.upper(), StreamARN=stream_arn )['DataEndpoint'] return boto3.client('kinesis-video-archived-media', endpoint_url=endpoint)
def get_hls_streaming_session_url(stream_processor_event: dict) -> str: # Generates an HLS streaming URL for the video stream. # See https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/kinesis-video-archived-media/client/get_hls_streaming_session_url.html
kinesis_video = stream_processor_event['InputInformation']['KinesisVideo'] stream_arn = kinesis_video['StreamArn'] kvs_am_client = get_kvs_am_client('get_hls_streaming_session_url', stream_arn) start_timestamp = datetime.fromtimestamp(kinesis_video['ServerTimestamp'], JST) end_timestamp = datetime.fromtimestamp(kinesis_video['ServerTimestamp'], JST) + timedelta(minutes=1)
return kvs_am_client.get_hls_streaming_session_url( StreamARN=stream_arn, PlaybackMode='ON_DEMAND', HLSFragmentSelector={ 'FragmentSelectorType': 'SERVER_TIMESTAMP', 'TimestampRange': { 'StartTimestamp': start_timestamp, 'EndTimestamp': end_timestamp, }, }, ContainerFormat='FRAGMENTED_MP4', Expires=300, )['HLSStreamingSessionURL']スタックのデプロイ
SAMアプリケーションをビルドしてデプロイします。
sam buildsam deploy顔のインデックス登録
USBカメラで顔を検出するには、事前にRekognitionの顔コレクションに顔をインデックス(登録)しておく必要があります。この目的にはIndexFaces APIを使用します。
以下を実際の値に置き換えてください。
<YOUR_BUCKET><YOUR_OBJECT><PERSON_ID>
aws rekognition index-faces \ --image '{"S3Object": {"Bucket": "<YOUR_BUCKET>", "Name": "<YOUR_OBJECT>"}}' \ --collection-id FaceCollection \ --external-image-id <PERSON_ID>Rekognitionは顔コレクションに実際の画像を保存しません。代わりに、顔の特徴をメタデータとして抽出し保存します。
https://docs.aws.amazon.com/rekognition/latest/dg/add-faces-to-collection-procedure.html
For each face detected, Amazon Rekognition extracts facial features and stores the feature information in a database. In addition, the command stores metadata for each face that’s detected in the specified face collection. Amazon Rekognition doesn’t store the actual image bytes.
ビデオプロデューサーのセットアップ
この例では、Ubuntu 23.10を実行するRaspberry Pi 4B(RAM 4GB)をビデオプロデューサーとして使用します。

GStreamerプラグインのビルド
AWSはAmazon Kinesis Video Streams CPP Producer, GStreamer Plugin and JNIを提供しています。このSDKにより、Raspberry PiからKinesis Video Streamsへの動画ストリーミングが可能になります。
AWSはGStreamerプラグイン用のDockerイメージを提供していますが、アーキテクチャの制約によりRaspberry Piでは動作しない場合があります。
以下のコマンドを実行します。システムのスペックによっては、ビルドに20分以上かかることがあります。
sudo apt updatesudo apt upgradesudo apt install \ make \ cmake \ build-essential \ m4 \ autoconf \ default-jdksudo apt install \ libssl-dev \ libcurl4-openssl-dev \ liblog4cplus-dev \ libgstreamer1.0-dev \ libgstreamer-plugins-base1.0-dev \ gstreamer1.0-plugins-base-apps \ gstreamer1.0-plugins-bad \ gstreamer1.0-plugins-good \ gstreamer1.0-plugins-ugly \ gstreamer1.0-tools
git clone https://github.com/awslabs/amazon-kinesis-video-streams-producer-sdk-cpp.gitmkdir -p amazon-kinesis-video-streams-producer-sdk-cpp/buildcd amazon-kinesis-video-streams-producer-sdk-cpp/build
sudo cmake .. -DBUILD_GSTREAMER_PLUGIN=ON -DBUILD_JNI=TRUEsudo makeビルドが完了したら、以下のコマンドで結果を確認します。
cd ~/amazon-kinesis-video-streams-producer-sdk-cppexport GST_PLUGIN_PATH=`pwd`/buildexport LD_LIBRARY_PATH=`pwd`/open-source/local/libgst-inspect-1.0 kvssink以下のような出力が表示されるはずです。
Factory Details: Rank primary + 10 (266) Long-name KVS Sink Klass Sink/Video/Network Description GStreamer AWS KVS plugin Author AWS KVS <[email protected]>...毎回環境変数をリセットしなくて済むように、~/.profileに以下のexport文を追加します。
echo "" >> ~/.profileecho "# GStreamer" >> ~/.profileecho "export GST_PLUGIN_PATH=$GST_PLUGIN_PATH" >> ~/.profileecho "export LD_LIBRARY_PATH=$LD_LIBRARY_PATH" >> ~/.profileGStreamerの実行
プラグインをビルドしたら、USBカメラをRaspberry Piに接続し、以下のコマンドを実行して動画データをKinesis Video Streamsにストリーミングします。
以下を実際の値に置き換えてください。
<KINESIS_VIDEO_STREAM_NAME><YOUR_ACCESS_KEY><YOUR_SECRET_KEY><YOUR_AWS_REGION>
動画品質を上げる(解像度やフレームレートを上げるなど)と、AWSのコストが増加する可能性があります。
gst-launch-1.0 -v v4l2src device=/dev/video0 \ ! videoconvert \ ! video/x-raw,format=I420,width=320,height=240,framerate=5/1 \ ! x264enc bframes=0 key-int-max=45 bitrate=500 tune=zerolatency \ ! video/x-h264,stream-format=avc,alignment=au \ ! kvssink stream-name=<KINESIS_VIDEO_STREAM_NAME> storage-size=128 access-key="<YOUR_ACCESS_KEY>" secret-key="<YOUR_SECRET_KEY>" aws-region="<YOUR_AWS_REGION>"Kinesis Video Streams管理コンソールに移動すると、ライブストリームを確認できます。

テスト
Rekognition Videoストリームプロセッサの起動
Rekognition Videoストリームプロセッサを起動します。このサービスはKinesis Video Streamを購読し、顔コレクションを使って顔を検出し、その結果をKinesis Data Streamにストリーミングします。
ストリームプロセッサを起動します。
aws rekognition start-stream-processor \ --name face-detector-rekognition-stream-processorストリームプロセッサが実行中であることを確認するため、ステータスを確認します。
aws rekognition describe-stream-processor \ --name face-detector-rekognition-stream-processor | grep "Status""Status": "RUNNING"と表示されれば正常です。
顔の検出
USBカメラが動画をキャプチャすると、Rekognition Videoストリームプロセッサが動画ストリームを解析し、顔コレクションに基づいて顔を検出します。
結果を確認するには、以下のコマンドでLambda関数のログを表示します。
sam logs -n Function \ --stack-name face-detector-using-kinesis-video-streams \ --tailログレコードには、以下の例のようにストリームプロセッサイベントの詳細情報が含まれます。
{ "InputInformation": { "KinesisVideo": { "StreamArn": "arn:aws:kinesisvideo:<AWS_REGION>:<AWS_ACCOUNT_ID>:stream/face-detector-kinesis-video-stream/xxxxxxxxxxxxx", "FragmentNumber": "91343852333181501717324262640137742175000164731", "ServerTimestamp": 1702208586.022, "ProducerTimestamp": 1702208585.699, "FrameOffsetInSeconds": 0.0, } }, "StreamProcessorInformation": {"Status": "RUNNING"}, "FaceSearchResponse": [ { "DetectedFace": { "BoundingBox": { "Height": 0.4744676, "Width": 0.29107505, "Left": 0.33036956, "Top": 0.19599175, }, "Confidence": 99.99677, "Landmarks": [ {"X": 0.41322955, "Y": 0.33761832, "Type": "eyeLeft"}, {"X": 0.54405355, "Y": 0.34024307, "Type": "eyeRight"}, {"X": 0.424819, "Y": 0.5417343, "Type": "mouthLeft"}, {"X": 0.5342691, "Y": 0.54362005, "Type": "mouthRight"}, {"X": 0.48934412, "Y": 0.43806323, "Type": "nose"}, ], "Pose": {"Pitch": 5.547308, "Roll": 0.85795176, "Yaw": 4.76913}, "Quality": {"Brightness": 57.938313, "Sharpness": 46.0298}, }, "MatchedFaces": [ { "Similarity": 99.986176, "Face": { "BoundingBox": { "Height": 0.417963, "Width": 0.406223, "Left": 0.28826, "Top": 0.242463, }, "FaceId": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", "Confidence": 99.996605, "ImageId": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", "ExternalImageId": "iwasa", }, } ], } ],}動画再生用のHLS URL
ログには、オンデマンド動画再生用に生成されたHLS URLも含まれます。例えば以下のようなものです。
https://x-xxxxxxxx.kinesisvideo.<AWS_REGION>.amazonaws.com/hls/v1/getHLSMasterPlaylist.m3u8?SessionToken=xxxxxxxxxxSafariやEdgeなど、対応しているブラウザでHLS URLを開いてください。
ChromeはHLS再生をネイティブでサポートしていません。Native HLS Playbackのようなサードパーティ拡張機能を使用できます。

クリーンアップ
作業が終わったら、ストリームプロセッサを停止しスタックを削除します。
aws rekognition stop-stream-processor \ --name face-detector-rekognition-stream-processorsam deleteまとめ
Raspberry PiのUSBカメラ映像をKinesis Video StreamsとRekognition Videoに流したところ、事前にインデックス登録した顔コレクションに対するリアルタイムの顔マッチが得られ、Lambda関数が各マッチを再生可能なHLS URLに変換しました。このパイプラインが独自の動画処理コードなしに機能するのは、Rekognition Videoのストリームプロセッサが、Kinesis Video Streamに対して直接顔の検出とマッチングを行っているためであり、Lambda関数には上記のとおり比較的シンプルな役割しか残っていません。とはいえ、kvssinkに渡すエンコード設定(解像度、フレームレート、ビットレート)はAWSのコストに直結するため、まずは本記事で使用した低い値から始め、検出精度が実際に必要とする場合にのみ引き上げる方が、最初から高品質をデフォルトにするより望ましいでしょう。
Related posts
Cognito User PoolsとOIDCでSlackサインインを実装する
Cognito user poolをOIDC経由でSlackと連携させ、"Sign in with Slack"をAmplifyでNext.jsアプリに組み込みます。
Lambda Web AdapterでFastAPIをAWS Lambdaにデプロイする
Lambda Web Adapterを使うと、FastAPIで書いたAPIバックエンドをコンテナのまま単一のLambda関数にデプロイできます。
API Gateway WebSocket:モック統合の実装
バックエンドのLambdaを一切使わず、モック統合のみでAPI Gateway WebSocket APIを構築し、あらかじめ用意されたレスポンスを返します。
CloudFront署名付きURL経由でS3にアップロードする
CloudFrontの署名付きURLを使えば、独自ドメイン経由でS3にアップロードできます。S3の直接の署名付きURLが使えない場合に有用です。
AWS EventBridge Scheduler:スケジュールに沿ってEC2を起動・停止する
Lambdaを介さずEventBridge SchedulerがEC2 APIを直接呼び出すことで、cronスケジュールに従ってEC2インスタンスを起動・停止します。
