I'm Ahmad Mustafa Anis, an AI researcher from Pakistan and master's student in AI Engineering at Carnegie Mellon University. Advised by Michael Tarr, I study self-supervised learning and computer vision.
Research shaped by vision, learning, and real-world systems.
I'm pursuing a Master of Science in Artificial Intelligence Engineering at Carnegie Mellon University. My work explores how visual models learn robust, transferable representations from large-scale unlabeled data.
I also lead the Geo-Regional Asia community at Cohere Labs, where I've hosted more than 50 researchers and helped connect AI practitioners across Pakistan and Asia.
02 Research
Learning useful structure from raw experience.
My work centers on self-supervised learning for visual representation learning, alongside a broader interest in world models.
01Self-Supervised LearningLearning robust visual representations from unlabeled data, with a focus on scalable methods that transfer across tasks and domains.
02World ModelsBuilding models that can understand an environment, plan useful actions, and reason about possible outcomes.
Self Supervised Contrastive Deep Learning in Computer Vision
December 2, 2023
In this session we learned how self-supervised learning is leveraged in Computer Vision by diving into SimCLR by Google, SimCLR v2, and CLIP by OpenAI.
Semantic search is a technique that uses deep learning algorithms to understand the context and meaning behind user queries, rather than just matching keywords. Attendees learned about the importance of semantic search, how it works, implementation in Python, and challenges.
Creating Interactive AI Applications: Deploying TensorFlow Models with Gradio and Hugging Face Spaces
February 26, 2023
Most models die in Jupyter Notebook and never reach a wider audience. With Gradio and Hugging Face Spaces, you can easily deploy your TensorFlow models with ease and create user-friendly interfaces for others to interact with.
Impact: Over 100 students, beginners, and industry professionals
Language Guided Recognition using CLIP (Machine Learning Focus Group)
December 31, 2022
Language Guided Computer Vision is now an active field of research in Computer Vision and surpases Supervised Learning. Learn how to use CLIP, a Language Guided Classifier by OpenAI by building a Semantic Image Search Engine.
How Convolutional Neural Networks work (Machine Learning Focus Group)
December 31, 2022
Convolutional Neural Networks are core-part of Computer Vision Deep Learning. They are widely used in real-world applications involving Object Classification, Object Detection, Object Segmentation, etc.
Convolutional Neural Networks are core-part of Computer Vision Deep Learning. This talk covers the history and theory of CNNs, their applications, and how easy it is to use it in Tensorflow 2.0.
Impact: Over 50 students, professionals and beginners
In-office session on how to use AWS rekognition for multiple tasks and on live-stream.
Impact: 10 Machine Learning Engineers
Webinars
"Data Science | Information, Trend, Road Map for Beginners"
April 16, 2023
An online webinar for University students discussing the important parts of Data Science and Machine Learning, roadmaps, general trends, tools and technologies, courses and how to get started with it.
Impact: 23 participants
Understanding SimCLRv2 by GoogleAI
April 12, 2023
An online webinar on SimCLRv2 and SimCLR by GoogleAI which showed the first time large scale application of self-supervised and semi-supervised learning on Images.
Impact: 15 participants
Language Guided Recognition | CLIP
November 19, 2022
An online webinar on Language Guided Recognition using CLIP by OpenAI which can perform zero shot classification and surpass the state of the art classifiers.
Impact: 20 participants
How Convolutional Neural Networks work
November 19, 2022
When it comes to Computer Vision, CNN plays a major role in creating an impact regarding how technology interacts with the world. From facial and handwritten recognitions to biometric authentications, CNN has been at the forefront of various breakthroughs within the Deep Learning space.
Hosted Talks at Cohere Labs Community (Selected only)
As a community lead at Cohere Labs Community, I've hosted over 50 sessions inviting guest speakers from Stanford, MIT, Google, UC Berkeley, Allen Institute of AI, and other top institutions to present their research. Below are some of the most popular talks:
Why you should learn Mathematics for Machine Learning
March 19, 2024
Talk by Cheng Soon Ong, senior principal research scientist at the Statistical Machine Learning Group, Data61, CSIRO, and author of "Mathematics for Machine Learning."
This paper investigates the image-level understanding of VLMs, specifically CLIP by OpenAI and SigLIP by Google. Our findings reveal that these models lack comprehension of multiple image-level augmentations.
Bridging the Data Provenance Gap Across Text, Speech and Video
S Longpre, N Singh, [8 authors], Ahmad Mustafa Anis, et al.