Start on Pinterest. Save the aesthetic you actually want, Y2K, old money, harajuku, clean girl, whatever it is, and take five references across to Arcads.
Five rather than one for a specific reason: AI softens an aesthetic every time it copies it. One reference gets interpreted, and the interpretation drifts toward generic. Five references from the same world hold the line, because the model has to satisfy all of them at once. You are not giving it a picture to copy, you are fencing in a style.
Generate at low resolution first and only upscale the one you like, so you are not paying full price to test.
A candid photo of [describe your person: age, hair, what they are wearing], [what they are doing], shot on a phone front camera in [describe the room and the light source].
Skin must look like real skin at close range: visible pores across the cheeks and nose, a little oil on the nose and forehead, faint redness under the eyes and around the nostrils, uneven natural tone, fine flyaway hairs at the hairline. Do not smooth, retouch or even out the skin.
Lighting is available light only, from [the window / the ceiling light], slightly uneven across the face, with a mild colour cast from the room. Not studio lighting, not a ring light, no fill.
The expression is mid-sentence and slightly asymmetric rather than posed or smiling at the camera. Framing is casual and a little off-centre, as if nobody composed it.
Match the style, grade and grain of the attached references.
This is where most people lose it. A perfect still turns into obviously-generated video the moment the motion is too smooth, so the flaws you write in here are camera flaws rather than skin flaws.
Animate this as a handheld phone video, roughly [8] seconds.
The camera is held in one hand: small constant drift and micro-shake, never a smooth glide, never a tripod. Once during the clip the autofocus hunts briefly before settling. When the subject turns toward the window, the exposure shifts and takes a moment to rebalance, blowing out slightly on that side.
The subject moves like a person mid-thought: small weight shifts, a blink that is not on a beat, one natural imperfect gesture. She does not perform to camera.
Keep the skin texture from the source image exactly as it is. Do not smooth, brighten or retouch anything as it moves.
No music, no transitions, no colour grading pass. It should look like an unedited clip straight off a phone.
Add any of these to either prompt. Reach for two or three, not all of them, because piling on every flaw at once reads as a filter rather than as reality.
SKIN AND FACE
visible pores across the cheeks and nose
oil on the nose and forehead catching the light
faint redness under the eyes and around the nostrils
a few fine flyaway hairs at the hairline
slightly chapped lips
uneven undertone, one cheek warmer than the other
an asymmetric expression caught mid-sentence
CAMERA
autofocus hunting once before it settles
exposure shifting when the subject turns toward the light
a slight fingerprint smudge softening one corner of the frame
mild rolling-shutter wobble on a quick movement
handheld micro-shake, never a smooth glide
framing slightly off-centre, as if nobody composed it
ROOM AND LIGHT
one practical light source with a colour cast
a slightly cluttered background that nobody tidied
mixed lighting, warm lamp against cool daylight
a shadow falling unevenly across the face
๐ฌ Arcads: arcads.ai