Skip to content

Why Google sees an empty page on your Base44 app, and how to fix it

Base44 hands crawlers a prerendered snapshot of each page, but only for routes written out as literal paths in App.jsx. Anything on a /:slug route gets a bare shell with your default title. How I caught it on my own site, the canonical tag that made it worse, and a two-curl test for yours.

Aatman (AJ) Jain8 min read
A magnifying glass hovers over a row of drawn terraced houses on mint paper, framing the one house that is only an empty dashed outline.
A magnifying glass hovers over a row of drawn terraced houses on mint paper, framing the one house that is only an empty dashed outline.

Base44 SEO mostly works out of the box: Base44 serves search engines a prerendered snapshot of each page, not the empty shell a React app starts as. The catch is that it only prerenders routes written out as literal paths in your App.jsx, so a page on a /:slug route, like a blog post or a case study, reaches crawlers as a bare shell with your default title and none of your content.

I found all of this on my own site. Here's how, and a two-minute test that tells you if your app has the same problem.

What a crawler actually gets from a Base44 app

A Base44 app is a single-page app. The server sends one small HTML file, then JavaScript builds the page in the browser. Base44's own docs (opens in a new tab) say it plainly: "Base44 apps are built to run in the browser, which normally means a crawler sees an empty page."

So Base44 does something smart about it. It "serves crawlers a rendered version of your app instead, with your meta tags, structured data, and real page content in the HTML". That's prerendering. Googlebot asks for a page, gets finished HTML, everyone's happy.

When a page hasn't been rendered, crawlers get the fallback instead: "a short summary carrying that page's heading, its description, and links to your other pages". Better than nothing. Not your blog post, though.

"But Google runs JavaScript?" It does, eventually. Google's JavaScript SEO guide (opens in a new tab) says a page waiting to be rendered "may stay on this queue for a few seconds, but it can take longer than that", and that "not all bots can run JavaScript." Base44's docs say the same about AI crawlers: they "read whatever version is ready at the time".

So yeah. You want the snapshot. Flip between the two views and you'll see why:

Interactive demo

Googlebot vs your browser, same URL

Flip between what Googlebot and a person get for the same page on a Base44 app. Base44 prerenders a snapshot for every literal <Route path> in App.jsx, so a crawler reads that page with its own title, h1 and canonical.

A page that only matches a /:slug route gets the app shell instead: no real content, just Base44's hidden summary (a generic title and h1 and a list of page links), not the page's own heading, text or canonical. A browser runs the JavaScript either way, so people never notice.

(If single-page app SEO is new to you, my SPA SEO glossary entry is the short version.)

The rule that caught me: literal routes only

Here's the bit nobody told me. Base44 works out which pages to prerender from the literal <Route path="..."> strings in src/App.jsx. A route with a parameter, like /work/:slug, isn't something it can list. It has NO idea which slugs exist. So every URL on that route gets the shell.

Literal routes get snapshots. Everything else gets the shell.

On 26 September I fetched my own site with a Googlebot user agent. The home page, /about and a handful of others came back as full snapshots. Lovely. Then I tried the case studies, a service page and a glossary term. Every one was the shell: the default title, a hidden summary with headings like "Gridiron Guide | Aatman Jain Portfolio", and a canonical tag pointing at my home page.

That last part is nasty, and I'll get to it. First, the routes.

The fix is boring, which is how I like my fixes. One literal route per page you know about, with the slug handed in:

jsx
{/* Before: one param route. Base44 can't list the slugs, so crawlers get the shell. */}
<Route path="/work/:slug" element={<CaseStudy />} />

{/* After: a literal route per case study. The param route stays as a fallback. */}
<Route path="/work/hattle" element={<WithParam name="slug" value="hattle"><CaseStudy /></WithParam>} />
<Route path="/work/gridiron-guide" element={<WithParam name="slug" value="gridiron-guide"><CaseStudy /></WithParam>} />
<Route path="/work/:slug" element={<CaseStudy />} />

WithParam is a small wrapper of mine that hands the page the same slug it would have read from the URL, so the page code didn't change at all. Passing the slug as a prop works too. Whatever you pick, the path has to be written out in full.

literal routes for pages that hid behind 3 param routes
38
the snapshot Googlebot gets for my security reviews page
202 KB
the shell it gets for a blog URL with no route
39 KB

Those 38 routes are 5 case studies, 6 services and 27 glossary terms. When I fetched my security reviews page as Googlebot on 28 September, I got a full snapshot with the page's own title, h1, canonical and structured data.

Then came the blog, with a twist. Posts are rows in my database, and some are scheduled. So each scheduled post gets its literal route ahead of time, and the page shows a noindex 404 until its date. I fetched this very post's URL as Googlebot the day before it went live: a full snapshot of my 404 page, marked noindex. Exactly right. My build now warns me if a live post has no route.

A made-up /blog/ URL still gets the shell, and it makes the difference easy to see. Here are both pages opened as Googlebot with JavaScript switched off. That's a rough picture of what a crawler that doesn't run JavaScript has to work with:

Literal route: the whole page is in the HTML
No literal route: an intro and a list of links, no post

The intro in the second picture is a <noscript> block I keep in index.html. Take it away and that page looks blank, because the fallback summary is squashed into a hidden 1 pixel box.

The canonical tag that pointed pages at my home page

While the site was being built, my index.html picked up a sensible-looking default:

html
<link rel="canonical" href="https://aatmanjain.com/" />

My React code then set the right canonical for each page as it rendered. The snapshots were fine, because Base44 captures them after that code runs. The shell was not. Every page served as a shell said, in its raw HTML, "the real version of me is the home page."

One line in index.html, copied into every shell

Google's guide says you "shouldn't use JavaScript to change the canonical URL to something else than the URL you specified as the canonical URL in the original HTML." Then it hands you the fix:

If you can't set the canonical URL in the HTML, then you can use JavaScript to set the canonical URL and leave it out of the original HTML.

Google Search Central, JavaScript SEO basics

So I deleted it, along with the matching og:url. It lived in my index.html for about twelve hours (26 to 27 September), long enough to show up in my bot check. Today my shell has no canonical at all, and each snapshot carries its own.

Your fallback summary is your app's name and description

That hidden summary leans on your app's settings: its name and its description. Mine were embarrassing. The app was still called "Aatman Jain Portfolio", and the description was an AI-generated line about a "digital sanctuary". That's what a crawler that doesn't run JavaScript read about me on any page without a snapshot.

I changed both through the Base44 API on 27 September. The name is now "Aatman (AJ) Jain", and the description says who I am and what I've shipped, in plain words. The same settings feed your link previews: Base44's docs say these previews "use the title, description, and logo you configure in your app's settings". It takes minutes and no code. Do it today.

The two-curl test

Here's how to check your own app. Fetch a page as Googlebot, then as a normal browser, and compare. Save this as two-curl.sh:

bash
URL="$1"
BOT="Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
HUMAN="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36"

curl -s -A "$BOT" "$URL" -o bot.html
curl -s -A "$HUMAN" "$URL" -o human.html

for f in bot.html human.html; do
  echo "$f: $(wc -c < $f) bytes"
  tr -s ' \n' ' ' < $f | grep -o '<title>[^<]*</title>'
done

This is my real output from the live site on 28 September. First a service page with a literal route, then a blog URL that doesn't exist:

Terminal
sh two-curl.sh https://aatmanjain.com/services/security-reviewsbot.html: 202425 bytes<title>Base44 security reviews and RLS checks · Aatman (AJ) Jain</title>human.html: 39530 bytes<title> Aatman (AJ) Jain · Developer and Base44 Partner in Belfast </title>sh two-curl.sh https://aatmanjain.com/blog/not-a-real-postbot.html: 39488 bytes<title> Aatman (AJ) Jain · Developer and Base44 Partner in Belfast </title>human.html: 39488 bytes<title> Aatman (AJ) Jain · Developer and Base44 Partner in Belfast </title>

Reading it is easy. If the bot gets a much bigger file with the page's own title, that's a snapshot. If both get about the same size and your default title, that's the shell. Notice the shell isn't empty: it's about 39 KB and carries the title from index.html. That's why this problem hides so well. Nothing looks broken.

Then open bot.html and look for the parts that matter. Here are the two files from that run:

Snapshot: own title, canonical, JSON-LD and h1
Shell: default title, hidden summary, no post

Check the canonical is the page's own URL and that your JSON-LD made it in. On my service page, both did.

Don't stop at the title

A snapshot existing isn't the same as a snapshot being right. A snapshot is a picture of your page at one moment. If that moment comes before the page has finished loading, it can catch your loading state instead of your content. The file size won't warn you, because a loading screen wrapped in your full layout can still be a big file.

So search bot.html for <h1 and read what's there. It should be the main heading a visitor sees on that page. No h1, or a loading message where your content should be? That page needs a closer look before you trust it.

Google's URL Inspection tool (opens in a new tab) in Search Console is the other half of this. Its live test lets you "view a screenshot of the rendered page as the Google-InspectionTool sees it", so you see the page the way Google's own tool rendered it.

One more trap: redeploying isn't rebuilding

My sitemap is generated at build time from the posts that are live, and it skips anything scheduled for the future. So when does a build actually happen? I tested it on 28 September:

Redeploy the same version

  • The build timestamp didn't move
  • So nothing was rebuilt

GitHub sync with new commits

  • The build timestamp changed
  • The sitemap and other generated files were rebuilt

If your sitemap or any other file is generated at build time, a new commit is what refreshes it. Publishing the same version again won't.

Checklist

  1. List your URLs

    Every page you care about: pages, case studies, posts, glossary terms.

  2. Write out the routes

    Give each page on a param route its own literal <Route path> in App.jsx. Keep the param route as a fallback.

  3. Drop the static canonical

    Remove any hard-coded canonical or og:url from index.html, and set them per page at runtime instead.

  4. Fix your fallback text

    Update your app's name and description in its settings.

  5. Run the two curls

    One URL of each type. Check the size, the title and the h1.

  6. Check again after deploys

    After a deploy that adds pages, run it again, and use URL Inspection's live test for the pages that matter most.

  7. Push to rebuild

    Push a new commit when generated files need refreshing.

If you'd rather someone else did the digging, this is the kind of thing I do in SEO and performance work. Or go run those two curls now. I'll wait.

FAQ

Does Base44 handle SEO automatically?

Mostly, yes. Base44 serves crawlers a prerendered snapshot of each page, with your meta tags and content in the HTML, and falls back to a short summary for pages it hasn't rendered. The gap is pages on parameter routes like /blog/:slug, which only get the snapshot if you add a literal route for each one.

Why do crawlers see my default title on some Base44 pages?

They're getting the plain app shell, which carries the title from your index.html. That happens when a page has no prerendered snapshot, usually because its route has a parameter such as :slug. Add a literal <Route path> for the page and check again with a Googlebot user agent.

If Googlebot runs JavaScript, why does the snapshot matter?

Google renders JavaScript in a second pass that can be delayed, and Google itself says not all bots can run JavaScript. A bot that can't only ever reads the raw HTML. A snapshot hands crawlers the real page on the very first request.

Do I need a custom domain for Base44 SEO?

Base44's docs say its SEO setup is "designed and supported for apps that are published on a custom domain". Free Base44 URLs can still be crawled, but they aren't meant for long-term SEO. Connect your own domain before you worry about snapshots.

Newsletter

Liked this one?

Get the next post in your inbox. Only when there is something worth sending.

No spam. Unsubscribe whenever.

Got an app in mind?

I build and fix Base44 apps, and I'm looking for a remote internship. Tell me what you're working on.