AI Image Caption Generator with HTML, CSS & JavaScript

40 DAYS 40 PROJECT CHALLENGE

Day #16

Project Overview

The AI Image Caption Generator is a beginner-friendly web application that uses HTML, CSS, and JavaScript. The project allows users to upload an image and generate a descriptive caption using an AI vision model. After users select an image, JavaScript reads the file, converts it into a Base64 data URL, and displays a preview before sending it for analysis. It also integrates an AI API to analyze the uploaded image. As a result, users can receive a dynamically generated caption directly on the webpage. This project provides practical experience with file handling, image previews, API integration, asynchronous JavaScript, and dynamic content rendering. It is useful for beginners who want to understand how AI-powered features can be added to a simple HTML, CSS, and JavaScript project.

Key Features

  • Image Upload: Users can select an image from their device using the file input. The input accepts image files so that the application can process visual content.
  • Image Preview: After selecting an image, JavaScript reads the file and displays a preview inside the application. This allows users to confirm the selected image before generating a caption.
  • AI Caption Generation: The application sends the uploaded image to an AI vision model along with a prompt asking the model to describe the image and generate a caption.
  • AI API Integration: The project connects to the OpenRouter API through a JavaScript fetch() request. The request sends the image and prompt to the selected AI model for analysis.
  • Base64 Image Conversion: JavaScript uses the FileReader API to convert the selected image into a Base64 data URL. This format allows the image data to be included in the API request.
  • Dynamic Caption Display: After receiving the API response, JavaScript extracts the generated caption and displays it inside the caption result section without reloading the webpage.
  • Upload Validation: Before generating a caption, the application checks whether an image has been uploaded. If no image is selected, it displays a message asking the user to upload an image first.
  • Loading Message: While the AI service is analyzing the image, the application displays an "Analyzing image..." message. This gives users immediate feedback while they wait for the API response.
  • Error Handling: The JavaScript code uses try...catch to handle API or request errors. If the caption cannot be generated, the application displays an error message instead of leaving the result area blank.
  • Responsive Interface: The layout uses a centered container with flexible sizing and spacing. In addition, the page includes padding so that the interface remains usable on smaller screens.
  • Clean User Interface: The project uses a simple card-based design with a heading, image upload field, preview area, action button, and caption result box. As a result, the main functionality remains easy to understand.
  • Dynamic Content Rendering: JavaScript updates the preview and caption sections directly in the DOM. Therefore, users can interact with the application without refreshing the page.

What You'll Learn

  • How to create an image upload interface using HTML.
  • How to access selected files using JavaScript.
  • How to use the FileReader API.
  • How to convert an image into a Base64 data URL.
  • How to display an uploaded image dynamically.
  • How to handle the change event for file inputs.
  • How to create asynchronous functions using async and await.
  • How to send requests to an external API using fetch().
  • How to display dynamically generated content on a webpage.
  • How to validate whether an image has been uploaded.
  • How to display loading and error messages.
  • How to use try...catch for API error handling.

HTML Code

HTML creates the basic structure of the AI Image Caption Generator. The main .container acts as the central section of the application and contains the project heading, image upload field, image preview, Generate AI Caption button, and caption result area, the #imageInput file input allows users to select an image from their device. It uses the image file restriction so that users can select supported image formats.

Next, the #preview image element is used to display the selected image. Initially, JavaScript keeps this preview hidden. Once an image is selected and successfully loaded, JavaScript sets its source which makes the preview visible and the Generate AI Caption button calls the generateCaption() function when the user clicks it. This function checks the selected image and then sends the image to the AI service for analysis.

Finally, the #captionResult element provides a dedicated area where JavaScript displays the generated caption, validation messages, loading messages, or error messages. The HTML file also loads the external script.js file, which controls the application’s interactive functionality.

<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<title>AI Image Caption Generator</title>
<link rel="stylesheet" href="style.css">
</head>

<body>

<div class="container">

<h1>AI Image Caption Generator</h1>

<input type="file" id="imageInput" accept="image/*">

<img id="preview">

<button onclick="generateCaption()">Generate AI Caption</button>

<div id="captionResult"></div>

</div>

<script src="script.js"></script>

</body>
</html>

CSS Code

CSS controls the visual appearance and layout of the AI Image Caption Generator. First, the universal selector removes the default margin and padding and applies box-sizing: border-box. It also sets Arial as the main font for the application. The body uses a dark blue background and Flexbox to center the application both horizontally and vertically. In addition, padding around the body provides some space between the application and the edges of smaller screens.

This .container creates the main white card that holds the application. It uses padding, rounded corners, a fixed width, centered text, and a subtle box shadow. Consequently, the project gets a clean card-based appearance and the image preview uses the #preview selector to occupy the available container width. It also has rounded corners and a top margin. Initially, the preview is hidden using display: none. JavaScript changes this property after an image is uploaded.

Finally, the #captionResult section provides a separate area for displaying the generated caption. It uses a light background, padding, rounded corners, and a minimum height so that the result remains visually separated from the rest of the interface.

*{
margin:0;
padding:0;
box-sizing:border-box;
font-family:Arial;
}

body{
background:#042354;
display:flex;
justify-content:center;
align-items:center;
height:100vh;
padding:20px;
}

.container{
background:white;
padding:30px;
border-radius:12px;
width:500px;
text-align:center;
box-shadow:0 10px 30px rgba(0,0,0,0.1);
}

h1{
margin-bottom:20px;
}

#preview{
width:100%;
margin-top:10px;
border-radius:8px;
display:none;
}

button{
margin-top:15px;
width:100%;
padding:12px;
background:#2563eb;
border:none;
color:white;
border-radius:6px;
cursor:pointer;
font-weight:bold;
}

button:hover{
background:#1d4ed8;
}

#captionResult{
margin-top:20px;
background:#f1f5f9;
padding:15px;
border-radius:6px;
min-height:80px;
}

Javascript Code

JavaScript controls the image upload, preview, AI caption generation, and error handling. First, it uses FileReader to convert the selected image into a Base64 data URL and displays it in the preview. Next, the generateCaption() function checks whether an image has been selected. If an image is available, it uses fetch() to send the image and prompt to the OpenRouter API. The function uses async and await to handle the API request.

After receiving the response, JavaScript extracts the generated caption and displays it on the page. Finally, a try...catch block handles API errors and shows an appropriate message to the user. JavaScript listens for changes to the file input. It then reads the selected image with FileReader and displays the image preview.

The generateCaption() function validates the image and sends it to the AI model using fetch(). The returned caption is then displayed dynamically. The API request uses try...catch to handle errors. If something goes wrong, JavaScript displays an error message instead of leaving the result section empty.

let imageBase64 = ""

const imageInput = document.getElementById("imageInput")
const preview = document.getElementById("preview")

imageInput.addEventListener("change", function(){

const file = this.files[0]

if(!file) return

const reader = new FileReader()

reader.onload = function(e){

imageBase64 = e.target.result

preview.src = imageBase64
preview.style.display = "block"

}

reader.readAsDataURL(file)

})

async function generateCaption(){

const result = document.getElementById("captionResult")

if(!imageBase64){
result.innerText = "Please upload an image first."
return
}

result.innerText = "Analyzing image..."

try{

const response = await fetch("https://openrouter.ai/api/v1/chat/completions",{

method:"POST",

headers:{
"Content-Type":"application/json",
"Authorization":"Bearer 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"
},

body:JSON.stringify({

model:"openai/gpt-4o-mini",

messages:[
{
role:"user",
content:[
{
type:"text",
text:"Describe this image and generate a caption."
},
{
type:"image_url",
image_url:{ url:imageBase64 }
}
]
}
]

})

})

const data = await response.json()

result.innerText = data.choices[0].message.content

}catch(error){

result.innerText = "Error generating caption."

}

}
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments

Related Projects

Day 13 : AI Blog Title Generator

Generates engaging blog titles based on user input keywords.

Concepts: User input handling, dynamic text generation, UI updates.

Day 14 : AI Code Explainer UI

Displays explanations for code snippets with a clean preview interface.

Concepts: Input handling, dynamic rendering, formatted output display.

Day 18 : AI Business Name Generator

Generates creative business names based on keywords and category.

Concepts: String manipulation, dynamic generation, UI interaction.