WipeMark

Sponsored

Advertise

How AI Background Removal Works: Masks, Matting and Inpainting

A model predicts a mask that says, per pixel, how much is subject. Here is how segmentation, matting, BiRefNet and LaMa inpainting actually do it.

By the WipeMark team9 min read

Key takeaways

  • Background removal predicts an alpha mask, a value from 0 to 1 for every pixel, and the cut-out is computed with the compositing equation C = αF + (1 − α)B.
  • Segmentation gives each pixel a hard yes or no, while matting estimates partial transparency, which is what hair, fur, smoke and motion blur need.
  • BiRefNet, published in 2024, is a dichotomous image segmentation model whose default weights work at 1024 x 1024 pixels and whose code is MIT-licensed.
  • Most tools run the model on a reduced copy and scale the mask back up, because model cost grows with pixel count and a 12-megapixel photo has about 11 times more pixels than a 1024 x 1024 input.
  • Object erasers use inpainting models such as LaMa, which uses fast Fourier convolutions to see the whole image and was trained on 256 x 256 crops yet works on much larger images.

AI background removal works by running a neural network that predicts a mask: a grayscale image the same shape as your photo, where white means "keep", black means "remove" and gray means "partly see-through". The tool then multiplies your photo by that mask to produce a cut-out with a transparent background. Everything else, from hair quality to speed, comes down to how good that mask is and at what resolution it was made.

This explainer covers the two kinds of masks (segmentation and matting), why hair and glass are hard, what models like U²-Net and BiRefNet do, why tools work at 1024 pixels and scale up, and how inpainting fills an area you erase.

The one equation behind every cut-out

Every cut-out, in every editor, follows one formula: the color you see equals alpha times the foreground plus one minus alpha times the background, written C = αF + (1 − α)B. Alpha (α) is a number from 0 to 1 for each pixel. At 1 the pixel is fully subject, at 0 fully background, and in between it is a blend.

The idea of storing alpha alongside color goes back to Thomas Porter and Tom Duff's 1984 SIGGRAPH paper Compositing Digital Images. Alvy Ray Smith and James Blinn's 1996 paper Blue Screen Matting framed the reverse problem, pulling F and α out of a photo, and showed why it is hard: for each pixel you know C but must solve for the foreground color, the background color and alpha at once. That is more unknowns than equations. A blue or green screen helps because it makes B known. A background remover has no such help, so it learns to guess.

Move the slider below to see what that means for one pixel at the edge of a strand of hair:

One edge pixel of blond hair, placed on a new background

Hair color Frgb(214, 178, 112)
Matte: αF + (1 − α)Brgb(101, 121, 150)
Hard mask (0 or 1)rgb(40, 90, 170)

With a matte this pixel becomes 35% hair and 65% studio blue. A hard mask rounds it to all background, which is why cut-outs from plain segmentation look jagged or helmet-like around hair.

A transparent PNG simply stores alpha as a fourth channel next to red, green and blue. For more on that, see transparent PNGs explained.

Segmentation vs matting

Segmentation decides, for every pixel, subject or not; matting decides how much. Both produce a mask, but they answer different questions.

SegmentationMatting
Output per pixel0 or 1 (or a probability that gets thresholded)Any value from 0 to 1
Good atSolid objects: products, furniture, cars, bodiesSoft edges: hair, fur, veils, smoke, motion blur
Typical failureJagged or "helmet" hair, halosSmudgy edges if the estimate is wrong
Classic inputsJust the imageThe image plus a trimap (older methods) or the image alone (newer ones)

In practice the line is blurry. Modern segmentation models output soft probabilities, and those soft edges behave a bit like a matte. Dedicated matting models go further and try to get the exact fraction right at every edge pixel.

Trimaps and alpha mattes

A trimap is a hint image with three regions: definitely subject (white), definitely background (black) and unknown (gray). A matting model only has to solve the gray band, which is a much easier problem than solving the whole photo.

PhotoTrimapAlpha matte
A photo, a trimap marking a gray unknown band around the hair, and the alpha matte a matting model produces from it. Mask colors are literal: black, gray and white.

The paper that made deep learning the standard for matting was Deep Image Matting by Ning Xu, Brian Price, Scott Cohen and Thomas Huang (CVPR 2017). It fed an image plus a trimap to an encoder-decoder network, refined the result with a second small network, and introduced a dataset of 49,300 training images and 1,000 test images.

Trimaps are tedious to draw, so later work removed the need for them. MODNet (Ke et al.) is one example: it does trimap-free portrait matting from a single image and reports 67 frames per second on a GTX 1080 Ti. Consumer tools almost always work this way. You upload a photo, and no one draws a trimap.

Why hair, glass and smoke are hard

These subjects are hard because their pixels genuinely contain both subject and background. A strand of hair is often thinner than one pixel, so each edge pixel is a mix: a little blond, a lot of sky. Three problems follow:

  1. Alpha has to be right. Snap it to 0 or 1 and you get jagged edges or a solid helmet shape. That is the difference the interactive above shows.
  2. The foreground color is contaminated. Even with correct alpha, the pixel's color still contains some of the old background, a "color spill". Placing the cut-out on a new background reveals a faint fringe of the old one, for example a green outline from grass.
  3. Glass is not just transparent. A wine glass bends and tints what is behind it. Alpha can only say "40% see-through"; it cannot store "and the background behind it is distorted". That is why glass cut-outs never look quite right on a new background, and why product photographers often shoot glass on the final background.

Smoke, veils, fur and motion blur all have the same issue as hair to a different degree. If you remove backgrounds from product shots, our guide to white-background product photos covers practical workarounds.

The models: U²-Net, DIS and BiRefNet

Most open background removers come from a line of research on salient object detection (find the main subject) and dichotomous image segmentation (cut it out very precisely). Three papers mark the path:

  • U²-Net (2020). Xuebin Qin and colleagues introduced a "nested U-structure" of ReSidual U-blocks, published in Pattern Recognition (arXiv:2005.09007). It was trained on the 10,553-image DUTS-TR dataset at 320 x 320 pixels. The full model is 176.3 MB and runs at 30 frames per second on a GTX 1080 Ti; a small version is 4.7 MB. It is still one of the models offered by the popular open-source rembg library.
  • DIS and IS-Net (2022). The same lead author then defined "dichotomous image segmentation", meaning highly accurate foreground-versus-background segmentation of objects with very fine structure, in Highly Accurate Dichotomous Image Segmentation (ECCV 2022). Its DIS5K dataset has 5,470 high-resolution images (2K, 4K or larger) with very fine-grained labels.
  • BiRefNet (2024). Peng Zheng and colleagues published Bilateral Reference for High-Resolution Dichotomous Image Segmentation in CAAI Artificial Intelligence Research. It splits the job in two: a localization module finds the object using global context, and a reconstruction module recovers fine edges using "bilateral reference", where image patches at several scales are the source reference and gradient maps are the target reference. Extra supervision on gradients pushes it to focus on fine detail.

The BiRefNet repository is MIT-licensed and offers several sets of weights: the default at 1024 x 1024, BiRefNet_HR trained at 2048 x 2048, a matting version, and BiRefNet_dynamic trained across resolutions from 256 x 256 to 2304 x 2304. Rembg lists several BiRefNet variants among its models as well. WipeMark uses BiRefNet too.

Resolution: why tools work at 1024 px and scale the mask up

Background removers usually run the model on a smaller copy of your photo and scale the resulting mask back up, because the cost of running a model grows with the number of pixels and the model was trained at a fixed size. BiRefNet's default weights expect 1024 x 1024. A 12-megapixel phone photo (4000 x 3000) has about 11 times as many pixels as that, and a model also tends to do best at the scale it was trained on.

Photo4000 pxCopy1024 pxMask1024 pxMask4000 pxmodel
The photo is shrunk for the model, the mask comes back at the small size, and it is scaled up to cut the full-size original.

Scaling the mask up works better than you might expect, for two reasons:

  1. Masks are smooth. Most of a mask is solid white or solid black, and those areas scale up perfectly. Only the edges carry detail.
  2. The colors still come from your original. The mask only decides how much of each pixel to keep. Every visible pixel in the cut-out is a pixel from the full-size photo, so the result keeps all its sharpness inside the subject.

The trade-off is at the edges. A strand of hair that is one pixel wide in the 1024 copy becomes about four pixels wide when scaled back to 4000, so the finest strands come out soft, and strands thinner than a pixel at 1024 may be lost. That is why high-resolution variants like BiRefNet_HR exist, and why a touch-up brush is still useful.

Here is how WipeMark applies this. A reduced copy, at most 1024 pixels, is sent to our server; BiRefNet returns the mask; the full-size cut-out is made in your browser; nothing is stored or logged. It usually takes about 20 seconds, and you get the full-size result with no watermark. Then you can fix any edge with the Restore and Erase brushes.

See the mask in action on your own photo. Free, full size, no sign-up.

Remove a background now

How inpainting fills an erased area

Inpainting is the reverse of cutting out: instead of removing the background around a subject, it removes an object and invents the background behind it. You paint a mask over the object, and a model predicts what the masked pixels should look like given everything around them.

The model behind many erasers, including WipeMark's, is LaMa (Resolution-robust Large Mask Inpainting with Fourier Convolutions, Suvorov et al., WACV 2022). Three ideas make it work:

  1. Fast Fourier convolutions. An ordinary convolution looks at a small neighborhood, so filling a big hole needs many layers before the model "sees" both sides. A fast Fourier convolution (Chi, Jiang and Mu, NeurIPS 2020) has a branch that works in the frequency domain, which gives it an image-wide receptive field in one step. That is why LaMa is good at periodic structures such as windows, fences, tiles and brickwork.
  2. Training on large masks. The authors trained on wide, aggressive masks, so the model learned to fill big holes rather than small scratches only.
  3. Resolution robustness. All LaMa models were trained on 256 x 256 crops, yet the repository reports that it "generalizes surprisingly well" to around 2,000 pixels. The largest version, Big LaMa, has 51 million parameters across 18 FFC residual blocks and was trained on 4.5 million images from the Places-Challenge dataset, according to the paper.

Diffusion-based fill tools take a different approach: they generate new content from a learned image distribution and can invent more complex structures, at a higher computing cost and with more risk of adding things that were not there. LaMa-style models are faster and more conservative, which suits removing small things from real photos.

WipeMark's eraser sends only the area around what you painted, at most 1024 pixels, to the server. LaMa fills it, and the patch is pasted back into your full-size photo in the browser. Because the fill uses only nearby context, small masks give the best results. Our guide on removing people from photos has a checker for whether a fill will work, and how to restore old photos applies the same model to scratches and date stamps.

What this means when you use a tool

Knowing how the model works tells you how to get a better result from any background remover or eraser:

  • Contrast helps segmentation. A subject that differs in color or brightness from its background gives the model clear evidence. Dark hair against a dark wall is the hardest case.
  • Resolution helps edges. Start from the original photo, not a screenshot. A bigger, sharper original keeps more real detail where the mask is soft.
  • Check hair, glass and the bottom edge at 100% zoom, and fix with a brush rather than starting over.
  • Pick formats that keep alpha. PNG and WebP store transparency; JPG does not.
  • Keep erase masks small and include shadows, so the fill has real context on every side.

For a hands-on walkthrough, see how to remove the background from a photo. If you are comparing services, our look at remove.bg alternatives covers what changes between them, most of it traceable to the model, the working resolution and the download size.

Questions people ask

How does AI know what the background is?

It does not know in a human sense. A neural network trained on thousands of photos with hand-made masks learns which visual patterns usually belong to the main subject. For each new photo it outputs a mask with a value per pixel, and anything near 0 is treated as background.

What is the difference between segmentation and matting?

Segmentation assigns each pixel to the subject or the background, a hard yes or no. Matting estimates how much of each pixel belongs to the subject, a value between 0 and 1. Matting matters at soft edges such as hair, fur, veils and motion blur, where one pixel really is part subject and part background.

Why is hair so hard for background removers?

A single strand of hair is often thinner than a pixel, so edge pixels contain a mix of hair color and background color. A hard mask must choose one or the other, which produces jagged or helmet-shaped edges. Good results need a soft mask and enough resolution to see the strands.

What model does WipeMark use for background removal?

WipeMark uses BiRefNet, an open model released under the MIT licence. It sends a reduced copy of the photo, at most 1024 pixels, to the server, the model returns a mask, and the full-size cut-out is made in your browser. Nothing is stored or logged.

How does an AI eraser fill in the area I remove?

It uses an inpainting model, which predicts plausible pixels for the hole from the surrounding image. LaMa, a widely used open model, relies on fast Fourier convolutions to take the whole image into account, which helps it continue repeating textures such as grass, brick and water. It invents rather than recovers, so unique details behind the object cannot come back.

Remove a background in seconds

Free, full size, no sign-up, no watermark. Nothing is stored.

Try WipeMark free

Keep reading