Hi, I'm Amey 👋

I build & deploy AI agents and LLM systems.

AI Engineer with an M.S. in Artificial Intelligence from University at Buffalo and production experience shipping LLM-powered systems — from AI agents backed by custom MCP tool servers to multimodal RAG pipelines and scalable Python backends on Azure and AWS.

01.

About Me

Hello! I'm Amey Managute, an AI Engineer with an M.S. in Artificial Intelligence from the University at Buffalo. I ship production LLM-powered systems — AI agents backed by custom MCP tool servers, multimodal RAG pipelines, and scalable Python backends.

I've built AI agents in Microsoft Copilot Studio with custom OAuth-secured MCP servers, multimodal image-generation engines on Google Gemini, and dual-index RAG pipelines with Pinecone and FastAPI. My work spans Azure and AWS, from stateless, horizontally scalable backends to containerized ML services.

I work most with Python, PyTorch, FastAPI, LangChain, and the major LLM APIs. Whether I'm architecting a retrieval pipeline or deploying a scalable service, I focus on building things that are fast, reliable, and maintainable.

Amey Managute
02.

Experience

AI Engineer

ProMazo (Client: Dynatrace)

Remote, USA Feb 2026 - Present
  • Built an AI agent in Microsoft Copilot Studio that automates document workflows across company research, Excel model computation (OpenPyXL), and PPTX rendering, powered by a custom OAuth 2.0-secured MCP tool server architected in Python, cutting analyst turnaround from hours to minutes
  • Engineered a stateless, horizontally scalable backend using Azure Redis for session state and Azure Files for artifact staging across replicas, with CI/CD via Azure DevOps and Azure Container Jobs automating TTL-based artifact cleanup to control production storage costs
Microsoft Copilot Studio MCP Python Azure Redis Azure DevOps

AI Engineer Intern

FutureHouse.ai

Remote, USA Feb 2026 - Jun 2026
  • Built a multimodal AI image-generation engine on Google Gemini Flash Image supporting up to 5 reference images with face identity preservation, background removal, and platform-aware output, resolving a production Postgres timeout by offloading assets to Supabase Storage
  • Engineered 4 LLM generation pipelines on Google Gemini behind a shared orchestration layer with multi-stage JSON repair
Google Gemini LLM Pipelines Image Generation Supabase

Machine Learning Intern

The Tann Mann Gaadi

Remote, India Jun 2023 - Aug 2023
  • Built a natural-language query interface for 50+ non-technical users, cutting average query time by 40%, and automated ETL and feature-engineering pipelines with LangChain and PandasAI, reducing data processing time by 60%
LangChain PandasAI Python Docker
03.

Education

Master of Science – Artificial Intelligence

University at Buffalo, The State University of New York

Buffalo, USA Aug 2024 - Feb 2026
  • GPA: 3.33/4.0
  • Specialization in Artificial Intelligence and Machine Learning
  • Coursework: Applied Machine Learning, Deep Learning, Natural Language Processing, Computer Vision
Machine Learning Deep Learning NLP Computer Vision

Bachelor of Engineering – Artificial Intelligence and Data Science

University of Mumbai

Mumbai, India Jul 2020 - Jun 2024
  • CGPA: 8.44/10
  • Focus on Computer Science fundamentals, Data Structures, Algorithms, and Software Engineering
  • Developed a strong foundation in programming and system design
Computer Science Data Structures Algorithms Software Engineering
04.

Projects

Multimodal RAG

A multimodal RAG pipeline that extracts text and figures from PDFs, using LLaMA-4 Scout to summarize figures, with a dual-index Pinecone design and intent-aware routing. Built with Jina CLIP v2 embeddings, a FastAPI backend, and S3-based ingestion.

Python FastAPI Pinecone Jina CLIP v2 LLaMA-4 S3

Manim MCP Server

A Model Context Protocol server that lets LLM clients generate 3Blue1Brown-style math animations from natural language. A 5-tool API cuts workflows from 30-50 granular operations down to 2-3 calls, with segment-based composition and a Pydantic-validated video rendering pipeline.

Python FastMCP Pydantic Manim

GPT from Scratch

Character-level generative language model built from scratch using GPT architecture. Includes tokenization, multi-head self-attention, positional encodings, and temperature-based text generation.

Python PyTorch

Fake News Detection

Built an RNN-based classifier on the WELFake dataset using PyTorch. Model Deployed with FastAPI and Streamlit interface achieving F1 score of ~0.80.

PyTorch FastAPI Streamlit

Image Forgery Detection

ELA + CNN pipeline to detect and highlight manipulated regions in images. Built with React frontend, Flask backend, and TensorFlow.

TensorFlow CNN React Flask
05.

Skills

Core Languages

Python C++ SQL

ML & AI Frameworks

PyTorch TensorFlow Scikit-learn Pandas NumPy

GenAI & LLMs

RAG LangChain Hugging Face Prompt Engineering MCP OpenAI API Anthropic API Groq

Backend & APIs

FastAPI REST APIs Pydantic Vector Databases (Pinecone) Redis

MLOps & DevOps

MLflow Docker LangSmith CI/CD Git Linux

Cloud (AWS & Azure)

S3 EC2 ECS Lambda Bedrock Azure Container Apps Azure Redis Azure DevOps
06.

Get In Touch

Let's Connect

I'm always open to conversations about AI, ML, and software development. Whether you have a question, want to collaborate, or just want to say hi, feel free to reach out!