AI Agent for Solution Engineering to Enable Factory Automation 

How AI Agents Autonomously Define Automation Opportunities

Solution Engineering Ensures the Right Solution Gets Deployed

Labor shortages are forcing many factories to consider deploying automation to reduce the ergonomically challenging manual work. Every factory automation engagement begins with understanding the customer’s current process end to end. This step is critical because deploying a robotic cell does not necessarily create value. A robotic cell can automate work and still fail to deliver business value to the customer if it does not deliver the right operational outcome. In these cases, the cell may be underutilized, left idle, or eventually decommissioned. Customers also do not always know what the right automation solution should be. They may approach a solution provider with a specific concept based on limited information about what is technically feasible or what will create the greatest operational impact. Solution engineering is the process that determines what should be automated, how it should be automated, and whether the resulting solution will deliver meaningful operational and business value.  It must look beyond the requested equipment and determine the underlying production needs and the operational outcomes the deployment must deliver.

Today, this work is performed by industrial engineers, solution engineers, and consultants who gather evidence from factory observations, process videos, standard operating procedures, time studies, layouts, production data, and operator interviews. They use this information to reconstruct the current process, quantify variability, identify production constraints, and define a solution that can deliver measurable value. This is especially difficult in high-mix manufacturing, where part variation, changing workflows, and differences between operators make the process harder to characterize and the resulting requirements less certain. As GrayMatter Robotics advances Factory Superintelligence, we are developing AI capabilities that make this reasoning more scalable and increasingly autonomous. As factories become increasingly operated and optimized by AI agents, the quality of this deployment intelligence will shape the robotic capabilities that are designed, commissioned, and ultimately brought into production.

Solution Engineering Requires Complex, Multi-domain Reasoning

Designing the right automation solution requires balancing several objectives at once. A manufacturer may want higher throughput, but the solution must also satisfy quality requirements, reduce rework, improve safety and ergonomics, fit within the available space, integrate with existing operations, and deliver a compelling economic return. These objectives are often interdependent and sometimes in conflict. Increasing automation may improve consistency, for example, but it may also introduce additional tooling, sensing, integration, or cycle-time constraints. Solution engineers must determine the right automation boundary within this trade space. They decide which tasks should be performed by robots, which should remain with people, and where human–robot teaming creates the best overall outcome. This requires reasoning about part geometry, process access, production variability, tool and sensor capabilities, cycle time, and the limits of the available robotic applications. In some cases, one robotic application can address the complete process. In others, the right solution may automate only the highest-value portion of the work or combine multiple applications over a phased deployment roadmap.

The resulting concept must then be translated into a defensible business case. Engineers estimate how the proposed solution will affect throughput, labor utilization, quality, consumables, rework, safety, and future production flexibility. They also determine whether the current product portfolio can deliver those outcomes with acceptable technical and commercial risk. The final recommendation therefore depends on more than identifying activities in a video or extracting information from documents. It requires connecting manufacturing operations, robotic capability, geometry, economics, and customer priorities into one coherent deployment decision.

General-purpose foundation models can accelerate parts of this workflow, including high-level video understanding, document interpretation, and information synthesis. But solution engineering cannot be reduced to prompting a pretrained model. Manufacturing processes have specialized vocabulary, physical constraints, geometric relationships, and value drivers that are poorly represented in general-purpose data. Delivering reliable recommendations therefore requires innovation beyond existing foundation models: manufacturing-specific representations, models grounded in robotic and process capabilities, and reasoning mechanisms that connect technical feasibility to customer value. The Solution Engineering Agent brings these capabilities together to build the deployment definition.

The Solution Engineering Agent Identifies What type of Automation Will Deliver Value

The Solution Engineering Agent brings together a collection of specialized engineering skills to transform factory observations into evidence-backed solution requirements that will guide the solution design and deployment. Each skill addresses a different part of the solution-engineering problem: understanding the current process, reconstructing the workspace, evaluating ergonomics, determining robot feasibility, estimating performance, and quantifying customer value. The agent coordinates these capabilities rather than relying on any single model to determine what should be deployed. Many of these skills can build on increasingly capable pretrained models. Depth estimation and 3D reconstruction can provide spatial understanding of a workcell; human-pose models can support ergonomic analysis; and vision-language models can provide high-level interpretation of videos, documents, and customer requirements. These capabilities substantially accelerate solution engineering. However, identifying and developing a valuable solution often requires reasoning at a level of manufacturing granularity that general-purpose models are not designed to provide. It matters not only that an operator is “sanding,” for example, but which region of the part is being processed, with which tool, for how long, under what geometric constraints, and whether that work can be transferred to a specific robotic capability.

GMR is developing manufacturing-specific process intelligence to bridge this gap. The Solution Engineering Agent uses models trained and evaluated with our proprietary manufacturing specific data (Refer to our article on FSI), together with GMR’s Foundation Model for Scene Understanding, to establish a common representation of the factory environment. The scene representation connects parts, tools, fixtures, people, and regions of the workspace with their geometry, relationships, and relevant affordances. This allows individual skills to reason from the same underlying context. A process-analysis skill can associate an observed action with the tool and part region involved; a spatial-analysis skill can evaluate access and coverage for that region; and a robot-feasibility skill can determine whether the required interaction is compatible with the capabilities and constraints of a GMR robotic cell. The same representation can carry uncertainty, allowing the agent to distinguish well-supported conclusions from assumptions that require additional customer information or engineering review. The Solution Engineering Agent composes the outputs of these skills into the artifacts required to make a deployment decision: the automation boundary, human-robot task allocation, cell concept, performance estimates, technical requirements, risks, and an economic value case. Each recommendation can remain connected to the evidence and assumptions that produced it, allowing engineers and customers to understand not only what is being proposed, but why. This deployment definition becomes the engineering foundation for downstream cell design, commissioning, and production. As these capabilities mature, solution engineering can progress from a largely manual, expert-driven process toward increasingly autonomous deployment intelligence, allowing manufacturers to identify and receive the right automation solutions faster while giving downstream agents a stronger basis from which to operate.

Case Study: Autonomous Process Analysis for Solution Engineering

We evaluated the Solution Engineering Agent on a customer factory-floor video representing approximately one hour of manual surface-finishing work. The footage contained short, interleaved activities such as sanding, inspection, consumable changes, and bondo application, along with ambiguous states such as an operator holding a tool without actively processing the part or working on small regions that may be difficult to automate. Traditionally, this analysis requires an experienced engineer to review the footage, classify each activity, estimate duration, and assess feasibility with GMR’s robotic capabilities. For a study of this density, the process can take several engineer-days.

The Solution Engineering Agent generated the time study directly from the footage, grounding each activity in the Foundation Model for Scene Understanding using information about the tool, part, and process region. Segments with insufficient visual evidence, such as those affected by occlusion, were flagged as uncertain rather than force-classified. Compared with human-generated ground truth, the activity timeline achieved 96% accuracy. On an hour long video, time-lapsed into a couple minutes, it takes a solution engineer 20 minutes to break the work down and then another 90 minutes for an industrial engineer to produce the complete time study, against under 3 minutes for the agent. At engagement scale, this compounds. A new customer engagement can generate roughly 2000 hours of footage and the time study consumes the majority of the 10 weeks between the initial touchpoint and a proposed solution to the customer. And because the agent parallelizes through the footage where engineers cannot, we can compress the longest stretch of the timeline, so the proposed solution and the deployment behind it can arrive weeks sooner. The time study is only one aspect of the agent’s output, it also generated a 3D representation of the workspace and extracted process- and part-level information for downstream reachability, coverage, and feasibility analysis, while identifying ergonomically demanding portions of the workflow where automation could provide additional value.

Authors

  • Sagar Jatin Joshi

    Sagar Joshi is a Robotics and AI Engineer at GrayMatter Robotics, where he works on systems that enable robots to understand, plan, and operate reliably in unstructured industrial environments. His work spans motion planning, learning, and multimodal AI, with a focus on bridging advances in AI into capable real world robotic systems. He holds an M.S. in Mechanical Engineering from the University of Southern California, where his research explored diffusion policies for vision, force, and language-conditioned robotic manipulation and disassembly, and has been published in IEEE Robotics and Automation Letters (RA-L) and IEEE CASE. Personal Website: https://sagarjoshi73249.github.io/

  • Omey Manyar

    Omey Manyar is a Lead Robotics Engineer at GrayMatter Robotics, where he works on Physical-AI systems that help robots operate autonomously in complex, real-world manufacturing environments. He holds a Ph.D. in Robotics from the University of Southern California, where his research explored how robots can learn to handle deformable objects using physics-informed learning. Along the way, he has had the privilege of working with teams at Toyota Research Institute, Amazon Robotics, and Rolls-Royce in Singapore. His research has been published at venues including ICRA, IROS, and ASME, and has been recognized with multiple Best Paper Awards. Personal Website: https://omey-manyar.com/

  • Satyandra K. Gupta

    Satyandra K. Gupta is Co-Founder and Chief Scientist at GrayMatter Robotics, where he leads the company's foundational research in physical AI, computational decision-making, and human-centered robotics. He holds the Smith International Professorship in the Viterbi School of Engineering at the University of Southern California and serves as the founding director of the USC Viterbi Center for Advanced Manufacturing. Dr. Gupta previously served as Program Director for the National Robotics Initiative at the National Science Foundation (2012–2014). He has authored more than 500 technical articles and delivered over 200 invited talks worldwide. He is a Fellow of AAAS, ASME, IEEE, NAI, SME, and the Solid Modeling Association. He serves on the Technical Advisory Committee for the Advanced Robotics for Manufacturing (ARM) Institute and a member of the Association for Advancing Automation (A3) Robotics Technology Strategy Board. In the past he served on the National Materials and Manufacturing Board. He is a former Editor-in-Chief of the ASME Journal of Computing and Information Science in Engineering. His research has been covered by the Economist, Forbes, LA Times, IEEE Spectrum, Smithsonian Magazine, and numerous other leading publications.

Share the Post: