A little detour · 2026.09.21

時間があったので、歌う動物のAI動画を真似してみた。

TikTokで見かけた動画が気になって、「こんな雰囲気のものを、自分でもつくってみたいな」と思いました。
そこから生まれたのが、ジェットコースターで歌うスカンクです。

画像・映像はAIで生成したものです。音を出してどうぞ。

今回使ったサイトと、それぞれの役割。

ChatGPT ↗

キャラクター画像づくりに。実際の制作ではこの会話の内蔵画像生成を使い、眼鏡やよだれかけを少しずつ調整しました。画像モデルの世代は、この制作記録では確定できていません。

fal|Kling AI Avatar v2 Standard ↗

採用版の動画生成に使ったサイトとモデル。画像・音声・動きの指示を渡します。ページの「Image」に画像、「Audio」に歌の音源、「Prompt」に指示文を入れる構成です。今回は日本語の歌に口を合わせることを優先しました。

fal|MiniMax H3 Max Turbo ↗

普段の日用品やコスメの商品紹介動画は、falのMiniMax H3 Turboでつくっています。この制作で使用したモデルの表記は「MiniMax H3 Max Turbo」。今回は、乗り込みから走り出す動きもこちらで試しました。こちらの試作では指定の日本語音源を入力していません。生成された音声を残した原版と、曲を差し替えた編集版を下で見比べられます。

BGMer ↗

日本語の音楽の配布元。ちゃもろめさんの「天命転折」を使いました。素材ごとの利用条件を確認してから使ってください。

生成は有料です。サービスの料金・入力項目は変更されることがあります。紹介内容は2026年9月21日時点。

主役は、ちょっとファンキーなスカンク。

猫もかわいいけれど、今回はスカンクに。丸いサングラス、ふわふわの耳当て、それから黄色い稲妻柄のよだれかけ。リアルな毛並みに、少しだけお茶目な小物を合わせました。

舞台はジェットコースター。「こんなところで、そんなに楽しそうに歌うの?」というギャップがあったら、おもしろいかなと思って。

動画のもとにしたスカンクのAI生成画像

まずは画像づくり。小物を少しずつ足していきました。

リアルなスカンクに、ミント色の耳当てと黄色いよだれかけ。サングラスは、最初のオレンジ色のものから、細い金色フレームの丸い形に変更しました。歌う口元を隠さないことも大事なポイントです。

丸い金色のサングラスに編集した段階のスカンク
丸い眼鏡に変更した段階の画像。このあと背景をぼかした画像を動画に使いました。

実際に使った、丸いサングラスへの編集指示

入力したのは、編集前のスカンク画像と、眼鏡の形の参考写真。下の指示は、そのときの実行記録です。参考写真そのものはここには掲載していません。

実際の画像編集プロンプトを開く
Precise edit. Image 1 is the seated photorealistic skunk wearing mint earmuffs, amber sunglasses and yellow purple-lightning bib on a red roller coaster; edit THIS image. Image 2 is a screenshot of a small dog wearing round sunglasses; use ONLY its sunglasses design as reference, do not copy dog, hands, UI or setting. Replace the skunk's amber plastic sunglasses with perfectly round dark charcoal-black lenses in delicate thin gold metal wire frames, a narrow curved gold bridge and thin temple arms, matching the reference dog's glasses. Naturally sized to the skunk's head, bridge rests on upper muzzle, temples recede into fur beneath earmuffs. Realistic lens reflections and contact shadows, subtle imperfect physical alignment, no pasted-on look. Keep muzzle and mouth completely unobstructed. Preserve everything else: same skunk, fur, mint fluffy earmuffs, acid-yellow fabric bib with purple lightning bolts and black-white checkered trim, paws on red lap bar, roller coaster scenery, daylight, vertical 9:16 composition. One full-frame realistic photo, no text, no collage.

最初からつくるなら、こんなプロンプト。

途中の編集をまとめて、一枚目からつくりやすい形にしたものです。こちらは共有用に整理した未実行のプロンプトで、掲載画像を一度で生成した指示ではありません。

再現用の画像プロンプトを開く
Create a photorealistic vertical 9:16 image of a real striped skunk sitting upright in the front seat of a roller coaster. A front-mounted camera faces the skunk at eye level. The skunk wears small perfectly round dark charcoal-black sunglasses with delicate thin gold metal wire frames and a narrow curved bridge, fluffy mint-green earmuffs, and a vivid acid-yellow cotton bib with purple lightning bolts and black-and-white checkerboard trim. The glasses physically rest on the upper muzzle, their thin temple arms disappear into fur beneath the earmuffs, with realistic contact shadows and subtle reflections of the sky and rails. The bib has natural folds, seams and shadows and sits below the chin. Its natural animal mouth and entire muzzle remain unobstructed. Black seat, red padded lap restraint, both small front paws resting on the restraint, fluffy striped tail beside the seat. Green trees, coaster tracks and a distant Ferris wheel in the background, natural daylight. Endearingly serious expression, realistic fur and animal anatomy. No human lips or teeth, no extra limbs, no text, no collage.

背景のぼかしは追加で調整しています。その編集指示の原文は今回参照した記録に残っていないため、実行済みの文章として再構成して掲載することはしていません。

かわいい画像ができても、動画は一筋縄ではいかず。

最初は、乗り込むところから始めるつもりでした。でも、場面が変わると急に別の映像になったり、走り出したと思ったら逆走したり。歌っているのに、ジェットコースターがほとんど動かないこともありました。

そこで今回は、乗車シーンを欲張らず、座って歌っている15秒に絞ることに。口の動きだけでなく、顔や肩の揺れも少し大きめにお願いしています。背景はぼかして、主役がちゃんと目に入るようにしました。

完璧なジェットコースター映像、とはまだ言えないけれど。何度か試した中では、この子がいちばん楽しそうに歌ってくれました。

使ったのは、画像と日本語の歌。

普段の商品紹介動画は、falのMiniMax H3 Turboでつくっています。今回は「用意した日本語の歌に、口を合わせたい」という目的があったので、主役の画像と音源をfalの「Kling AI Avatar v2 Standard」に渡す方法にしました。音楽はBGMerで配布されている、ちゃもろめさんの「天命転折」。その一部分を使っています。

別のモデルでは、モデル自身がつくった違う曲で歌う結果にもなりました。今回は日本語の歌に合わせたかったので、用意した音源を入力できる方法を採用しています。あとから曲だけを差し替えても、口の動きはその歌に合ってくれないんですよね。

うまくいかなかった動画も、並べてみます。

「前へ進んで」と書けば、そのとおりになると思っていたのですが…。途中の動画を見ると、どこで悩んだのか分かりやすいかもしれません。いずれも採用前の試作です。

01|歌っているけれど、まだホーム?

乗り込むところから歌い始めるまでを、1回の生成でつなごうとした20秒版です。途中でホームを確認するような仕草はかわいい。でも、ホーム付近にいる時間が長く、期待した加速には届きませんでした。終盤には動き出す様子もあるので、「まったく発車しない」わけではありません。

モデル:fal-ai/kling-video/ai-avatar/v2/standard / 入力画像:boarding-v6.png / 音声:boarding-song-20s-v12.mp3

この試作で実際に使ったプロンプト
ONE UNBROKEN 20 SECOND TAKE, with continuous physical action from the supplied image. The same realistic small skunk wearing thin gold round dark glasses, mint earmuffs and yellow lightning bib climbs into the red roller coaster car. 0-3 seconds: climb in from the platform, turn to face the front of the vehicle and sit. The camera smoothly moves with the skunk into a position attached to the front of the car, looking back toward its face. 3-5 seconds: lower the red restraint and grip it with both paws. Once seated the camera stays rigidly fixed to this car for the ENTIRE remaining take. At 5 seconds the car starts rolling gently FORWARD, in the direction the skunk faces, toward the camera mount. Between 5 and 10 seconds speed gradually increases. The station BEHIND the skunk becomes progressively SMALLER, shrinking toward the background vanishing point. Station supports and trees travel from the near side edges inward toward the distant center BEHIND the skunk; they recede, never approach the camera. Between 10 and 20 seconds continue rolling FORWARD steadily along a shallow descending track and a gentle right bend. No backing up, no stop, no bounce, no backward launch. Keep the car and skunk facing their direction of travel. The track already traveled remains behind the seat. No looping or spinning. During the first five silent seconds the skunk does not sing. When the supplied singing starts, sing exuberantly directly toward the camera with wide open jaw movements aligned to the audio, bob the head and sway shoulders while safely seated. Fur and bib gradually flutter backwards toward the seat as headwind increases. Soft photographic background bokeh, but the receding station and trees remain recognizable enough to show travel direction. Preserve the same character, accessories, lighting and vehicle throughout. No cuts, no dissolves, no sudden change of framing, no teleport, no second shot, no scene transition, no text. Continuous boarding and forward departure are the priority.

02|あと5秒足したら、今度は逆走。

ひとつ前の動画の最後のフレームを画像として渡し、続きの5秒を別に生成しました。元動画の時間をそのまま延ばす機能ではありません。「ホームが遠ざかる」と具体的に書いたものの、確認すると逆走して見える結果に。下は追加部分だけの動画です。

モデル:fal-ai/kling-video/ai-avatar/v2/standard / 入力画像:continuous-v12-last.png / 音声:singing-continuation-5s-v14.mp3

この試作で実際に使ったプロンプト
Continue this exact shot for five seconds, beginning with the exact supplied final frame of an existing video. This skunk has just checked the loading platform and the coaster is NOW LEAVING. Preserve the same camera angle, framing, seat, red car, accessories and skunk scale at the start. This is the continuous forward departure, not another boarding scene. Car accelerates gently FORWARD, towards the direction in which the seated skunk faces, leaving the wooden station behind. The visible wooden platform and posts recede rapidly behind the car and shrink into the distant background; the station exits the frame behind. Background changes continuously from station to trees beside the outgoing track. Never return toward the platform, never reverse direction or stop. Camera travels with the car and gradually centers on the skunk as it turns its attention from the platform back to the lens, smiling and singing the supplied song with large expressive jaw movements precisely aligned to the audio. Cheerful head bobbing and shoulder rhythm; paws stay on the red lap bar. Increasing headwind ruffles fur toward the back of the seat, mint earmuffs and yellow lightning bib flutter. Thin gold round dark glasses remain on face. Soft optical background bokeh, natural sun and realistic fur. No scene cut, dissolve, teleport, unrelated scenery, camera zoom, loops or text. Prioritize physically continuous FORWARD acceleration and leaving the station while maintaining singing.

03|映像はよさそう。でも、日本語の歌が合わない。

MiniMaxの試作では、元から生成された曲には口の動きが合っているように感じました。ただ、今回使いたかった日本語の歌とは違う曲。あとからBGMerの音楽に差し替えてみても、口の動きまで変わるわけではありませんでした。

生成された音声を残した原版

日本語の歌に差し替えた編集版

原版は15秒。編集版は映像を20秒に引き延ばし、冒頭4秒を無音にして日本語の歌を載せています。比較条件が同じではなく、音声に合わせた口の再生成も行っていません。

MiniMaxの原版生成で実際に使ったプロンプト
A single uninterrupted photographic 15-second shot of this exact skunk boarding this red roller coaster and DEPARTING the station. Continuous physical motion. 0-3 seconds the skunk climbs from the wooden platform into the seat, sits facing the front of the car, lowers its lap bar. The camera smoothly tracks around from the platform to a front-mounted selfie camera looking back at the passenger. 3-7 seconds the car rolls forward, leaving the station gradually, smoothly picking up speed. The wooden station platform and roof supports recede behind the skunk and disappear into the distance. 7-15 seconds the car travels forward along the track into a tree-lined descending curve. The empty loading platform is now far away. Show continuous clear travel out of the station into a DIFFERENT outdoor stretch of track, not the same station scenery. Near trackside posts pass out behind the car while their size shrinks into the far background. Camera travels at exactly the same speed as the vehicle and looks back at the forward-facing seated skunk. Car never rolls backwards, never stops, never returns to loading platform. Skunk beams at the lens and sings joyfully with large jaw openings, head bobbing and shoulders rocking from 4 seconds onward. Both paws safely grip the red bar. Fur streams toward the seat behind its head as speed rises; yellow lightning bib flutters. Same gold round glasses and mint earmuffs, same skunk and car, photorealistic fur and natural lighting. Soft background bokeh with recognizable motion, no camera cuts or transitions, no static talking portrait, no full inversion or loop, no text.

ここから、乗り込みをいったん諦めて「最初から乗って歌う15秒」に絞りました。これは今回の試作での判断で、モデル全体が乗車シーンや歌唱に対応できない、という結論ではありません。

15秒の生成には、いくらかかった?

項目費用の目安
採用版と同じ設定の動画生成1回約0.85ドル
計算約15.2秒 × 0.0562ドル/秒

1ドル150円で換算すると、1回およそ128円。これは動画生成1回分の目安です。途中の試作分やChatGPTの契約料金は別なので、「全部で128円で完成」というわけではありません。

2026年9月21日時点の単価に基づく目安。同設定の過去の生成で0.85424ドルの請求を確認しています。採用版の確定請求額を示すものではありません。最新料金はfalのモデルページをご確認ください。

今回、動画に渡した指示。

参考になりそうなので、実際に使った英語のプロンプトも残しておきます。同じ指示でも、毎回まったく同じ映像になるわけではありません。

The same skunk enthusiastically SINGS directly TO THE CAMERA, accurately articulating the supplied Japanese song with BIG, exuberant singing gestures. The lower jaw drops visibly WIDE on open vowels, stays wide through sustained sung notes, and closes clearly with the syllables; keep exact audio timing and natural skunk mouth anatomy. Make the articulation large and readable, not tiny lip movements. The skunk enthusiastically bobs its head up and down and sways it left and right to the beat, shoulders rocking with the head, like it is joyfully belting out the song. Keep the face oriented toward the lens and the mouth visible throughout. Both paws hold the red safety bar. Preserve thin gold round sunglasses, mint earmuffs and yellow lightning bib. The roller coaster is ALREADY MOVING FORWARD at the opening frame and continues forward smoothly for the entire 15 seconds on a plausible course of your choice. The skunk faces the direction of travel; the camera is mounted in front of its face and travels with the same car. Everything visible behind the skunk is scenery the car has ALREADY PASSED: the red track hill and distant wheel recede and become smaller as the car moves farther away from them. New scenery appears alongside then recedes behind. Never move backward toward those landmarks, never stop, never return to the starting location. Keep the camera fixed to the car, with no camera orbit or zoom. A STRONG CONTINUOUS HEADWIND blows from the camera toward the skunk and its seat. The long cheek fur, white crown stripe and fluffy tail visibly stream BACKWARD and ripple continuously in the wind, individual tufts repeatedly bending and lifting. Mint earmuff fluff trembles and the loose bib edge flaps against the chest. This wind animation remains clearly visible throughout, while the sunglasses and earmuffs stay securely attached. Do not substitute camera shake for head bobbing or windblown fur. Sharp face and mouth, soft photographic background bokeh. One continuous shot, realistic skunk anatomy, no human lips, no cuts, no boarding or arrival scene. Preserve the supplied Japanese song and its timing exactly.

少し寄り道してみたら、かわいい子ができました。

思ったとおりに動いてくれなくて悩んだ分、ちゃんと歌ってくれたときはうれしい。日用品やコスメの紹介とは少し違うけれど、こういう「気になったから、やってみた」も、ときどき楽しんでいけたらと思います。

音楽:ちゃもろめ「天命転折」/BGMer
配布元の利用規約に沿って使用しています。動画の画像・映像はAIで生成しています。
← 読みものに戻る