Summary
This episode focused on FinAI and managing costs across AI program lifecycle stages. Isaac hosted enterprise architect Rameshwar as the special guest to discuss cost management from innovation through production and agent lifecycle phases. The discussion covered how token prices have dropped 90% over three years while enterprise AI bills have doubled, emphasizing that businesses need to look beyond simple tokenomics to total cost of ownership (TCO) which, includes infrastructure, data costs, and security measures. Key points included the importance of proper governance implementation from the start rather than adding it after deployment, the need for education on appropriate model usage and cost limits, and the challenge of measuring return on data when only a small percentage of stored data is actually analyzed. The panelists discussed chargeback models for different departments, the 11-layer architecture complexity in building effective AI agents, and the critical need for cybersecurity insurance coverage that properly accounts for AI-related risks.

Speakers
- Host – Isaac Sacolick
- Guest speaker – Rameshwar Balanagu, Executive Leader at the Intersection of AI, Cloud, Quantum & Human Insight
- Digital Trailblazers – Derrick Butts, Martin Davis, Joanne Friedman, John Patrick Luethe, Liz Martinez, Heather May, Joseph Puglisi, Elena Putilina
Discussion
- Innovation: Token prices dropped over 90% in three years, and enterprise bills doubled anyway. How should businesses manage AI spend when it’s hard to project and measure costs?
- Production: Compare AI to the cloud migration wave. What did we learn about cost then that we’re clearly not applying now?
- Lifecycle: Deployed AI agents are a permanent operating expense. What disciplines are needed to avoid AI debt, and what happened with apps that lacked ongoing support? Do AI agents need a chargeback model?
Research
Whiteboard

Transcript
[00:00:01] Speaker A: Welcome to our 181st episode of the Coffee with Digital Trailblazers. Every week we handle and discuss another topic of interest for digital transformation and AI leaders.
And really excited to be talking about a subject today that came up a few times during previous coffee hours and a number of people asked us to cover it. We’re going to talk about FIN AI, the finances behind investing in your AI capabilities. It’s a little bit of a riff of finops.
Oh gosh, look what I forgot to do.
I forgot to mute my own browser window. You guys could all laugh on the background around that.
I am back. Anyway, we’re talking about FINAI today and a little bit of a riff around finops and talking specifically around costs and in your AI program and how we manage them.
We will look at this subject from a number of different vantage points. I want to talk about innovation first, what it takes to experiment, what it takes to do our POCs and how we think about managing the costs around that particular life cycle. We’ll then talk and shift gears and talk about production and what it takes to bring an AI experiment into production, how we think about the average costs around that. And then we’ll get into life cycle what happens when it’s in production, Adoption’s increasing, cost changes, models change.
We have to think about model drift and things like that.
And really excited to be able to talk about all three life cycles around that. Folks, I’m going to paste a URL into the chat window. Do say hello in the chat window. I have opened up a poll today around upcoming coffee hour topics. It should take you five seconds to open this up and look at the four topics I am suggesting. This will likely be topics of choice going into September or October and I have my topics laid out for August. I’ll start announcing them later, but I do hope you can all just click into that link and tell us what you want us to be talking about. And if I missed any, do leave a comment in the comment stream of that post or on the comment stream here or just get back to me and let me know what you would like us to cover. So anyway, that’s the poll today. So let’s get to the research slide and I’m gonna share with you Some information once LinkedIn Zoom lets me do it. Here we go.
And Zoom is acting up today.
Hold on.
Zoom is not letting me do anything.
Oh, here we go. Hopefully you can all see that we’re talking about.
FIN AI and managing costs. Innovation is giving you a sense of some of the data Points I was able to find around AI and costs around it. And one of the things I found is just, wow, can I get a vote from somebody in my program? Can you hear me?
[00:03:50] Speaker B: Yeah. Yeah. Okay, but you. It stammers a little bit, Isaac, and
[00:03:56] Speaker A: it’s not live yet, right?
Well, my whiteboard just crashed.
The zoom whiteboard just crashed. You can believe that of all the things to go wrong and the last time this happened, it came back up.
But let’s, you know what, let’s pivot and let’s just jump into our conversation and maybe if my whiteboard does come back, I’ll be able to bring that back up and I’ll be able to read through my research slide. If you want to see the research slide, it’s@drive.starcio.com Coffee it’s listed. You can look at it there. The left side is talking about some data points around costs and the right hand side is talking about the different areas of focus around cost, from innovation through production and life cycle. I want to welcome our guest speaker today. Ram. Welcome to the group.
Ram is a longtime enterprise architect, a friend of the Coffee Hour. He’s been in the common stream many, many times.
A good friend of mine through our, our friendship together and today, Rahm, I want to talk about innovation first. Okay. And the very first thing is it’s, it’s sort of paradoxical token prices have dropped over 90% in the last three years, but our bills have died doubled and most CIOs and architects have become a lot more cost conscious because of watching these bills come in.
You’ve been very hands on in these environments. How should businesses manage the spend on AI innovation and when there’s challenges and how do we deal with the challenges around projecting costs and, and measuring costs? Ram, welcome to the group.
[00:05:56] Speaker C: Thanks, Isaac. And thanks to many of the other guests. It’s an honor for me to be on this side.
Let’s not put a damper. Let’s talk about innovation and excitement. Just how AI is doing some cool things. I was walking yesterday night, 10pm and I think I heard a dentist, she was telling her fellow team member, use AI for, for radiology and use it to build an SOP now. Didn’t get into what she was speaking, but I was like, what? This is at 10pm look at the power of what AI is doing now. That is from a consumer perspective. Then I was just talking to a friend of mine back in Europe.
They’ve had some issues.
We all know cable. We all know what is happening with mitos and all they in less than four hours they could reconstruct an SOP book. They were able to find critical vulnerabilities, they were able to simulate them and they were able to fix them.
This is the power of AI that we are living in. It’s matter of hours, not days, weeks, months and again we can talk and all now why is it that there is so much excitement and aha moment that we have? Even I myself everybody is now using AI like the hell out of it, but let’s peel the skin off it. When the radiologist or this dentist said you know, let’s use AI upload as X ray. The first thing that started flashing my bulbs is last week Satya Nadhala made a statement and you if you haven’t read it, it’s a reverse information paradox. What’s happening? So we will get into the tokenomics but what we are not accounting is we are paying.
Are we getting robbed out of our ip?
We are paying for the token cost but by training or giving the data the biggest mode any company has data.
We are using the frontier models for everything at the purpose of whom. So there is a direct material cost and indirect material cost which you cannot even multiply. That you know, I keep telling my friends is how do we show like if you have read this famous book on infonomics by Doug Lenny, direct monetization and indirect monetization you can show the cost of a token, you can show the cost of an infrastructure, but what about the cost of the data?
How do you show that in the tco? I don’t think we have an answer but what we can do is regulate how the architecture ends, how the tokenomics are and moreover from a finops. So this is something that blew my mind last week office I was always scared. But if this statement comes from Satya Nadala we all have to be very clear and cautious as we are looking at the tco. We need to look at an outcome. We need to look at an outcome but we also need to make sure that the outcome has some ways to protect our data. Because that line item is not shown on your balance sheet. You may be $100 billion company but the frontier models, if they have your data, don’t want to name it but you have seen Figma versus Cloud Design, $118 billion company being acquired by Adobe and now cloud design is much more powerful and you can’t do tokenomics on those kind of things.
[00:09:10] Speaker A: So it’s really interesting that you bring up value and data because it’s a frequent topic here at the coffee with digital trailblazers. And you know, usually what one of my guests will say, let’s focus on value first and then let’s see, what’s the word?
What’s the level of investment that’s worthy for this exploration that we’re going to go do. And so we set a budget, right? We think this is a billion dollar idea. We’re going to put a budget of 10 million for our POC and so that now we’re starting to, you know, cap our costs a little bit.
Do you have any suggestions on how we measure the cost when we start getting into using models? What are some of the techniques you’re using?
[00:09:58] Speaker C: You know, coming from architecture, I always go 30,000ft. But I’ll get to your answer just before.
Sorry, go ahead, Isaac.
[00:10:05] Speaker A: No, go for it.
[00:10:06] Speaker C: Derek and Joanne had their hands raised.
[00:10:09] Speaker A: So it’s a follow up. I’ll get to Derek and Joanne.
[00:10:13] Speaker C: Yeah, so before I answer your question, I want to give something Right before our call, I was just googling Flexera. What is the state of finops and what are the challenges? You won’t believe it’s been 10 years if you’ve been in. I did my certification in Aptio in 2018, so it’s 10 years roughly or you know, eight years. But finops, the word finops we used to, if you remember long back, we used to call this word called it finance management. It still is there, but it transformed into more glamorous. XVI called FinOps. 29% of the cloud computing is still idle.
That blew me out. The second thing is getting visibility into cloud is still a challenge. And the third thing is security.
What these three are still telling us, we’re still struggling with finops. Straight out, straight out. Now we didn’t talk till January to economics. We celebrated tokens like there’s no, there’s never, you know, use it. Celebrate leaderboards. I like to use this metaphors. You can shoot on me. But Jensen Wong said that if, if I pay an engineer of 500,000 salary, I expect him to use 250k. Soon every company started celebrating leaderboard jamboards.
What is the capital of France?
Do I need a frontier model? This is how we started becoming. We started using AI sdlc for everything. We forgot Google or we forgot our natural intelligence. We started using it, then git came, gave a sticker price. If you’re using it for your AI SDLC, we are going to do a 3x multiple. So the $20 cushion was never enough P&L and we were happily looking at our standard cloud models and infrastructure models. Now sudden started thinking of models are going down. Models are commodity. Kimi came open, weights came. But the problem was with the harness engineering and the loop engineering.
This is where people started thinking of you know guys or folks we need to start thinking and I would. I hate this word tokenomics for a reason Isaac. I googled multiple times.
Tokenomics is a basic simple math of input tokens, output tokens, times your discount price or whatever you work with your vendor. But the problem with that is tokenomics is just giving you a glimpse of what it is. There are hundred layers below the iceberg that are needed to give you a true tco. If if you ask me from an architecture tokenomics is not the true picture. FinOps is your TCO. And when we are talking outcome based again I will contradict myself. You know coming from my architecture we have pattern anti patterns.
What do you call an outcome. So if I’m able to increase my revenue, you know we say you know for simple thing not naming any of my companies or in the past if I could improve the revenue year over year or quarter over quarter around 6 to 10%. It’s phenomenal. My EBITDA has increased but two quarters down the lane. And the fact is AI goes rogue and somehow we forgot to add governance. There is a recall or a lawsuit. Now what happens to my outcome and the lawsuit is more expensive than what AI has produced.
Where did I justify my tco? Where did I justify my outcome? The classic case is OpenAI. We all know this Monday if OpenAI was not able to control its agents then they have a governance and security issue. So if we are talking tco it’s funny enough why I bring this is we are only looking at the angle we want to show.
And what I mean by that is if you see the finops.
I don’t know how many of you are very comfortable with the TBM unified method. Everything tries to a business layer but it has pools and sub pools and towers and connecting that dots to Flexera where it says more than 60 to 70% are still struggling with management.
The reason is you are only looking at an infrastructure like not. You know Isaac, you told me very beginning at the upfront you are an architect so go deep. If you see the model tokenomics of how it is we just said input output and whatever is caching and all but now you should start peeling the skin of the onion because we want to go deep. When you open a browser or your phone Open an app. Even in an enterprise, you know, whether using a Chat LLM or a copilot or whatnot, you are touching OSI layer.
When you open a browser, your browser is being secured with an oauth and it’s being connected to a cdn. And when your Chat LLM is working in. I’m sorry, why do I call Chat LLM? Sorry, it’s a copilot or any of them. These are authorized by you at that point. There is a runtime monitoring being done even before you opened any prompt. There is an edr, mdr, xdr, a sim tool doing even before you snoozed. There are six to seven tools before you even put a prompt.
This is a kind of invisible cost. And then you put your chat saying what is the best architecture design or how do I improve my top line using my inventory?
When you put the simple question of how do I improve my top line? And I want to reduce my inventory, I want more visibility into it. Three simple questions, very painless, very harmless. What is my inventory? How do I get more inventory visibility and how do I improve my top line?
Three questions, no harm, no pun intended. But if you see here when these three questions were asked, now you have, you know, when these three questions are put, the first thing is these three questions are being inspected. Is, is there anything like any dangerous data being asked? Is Isaac the right person to ask? You know, luckily in this case it’s inventory. But imagine if you say I want to know Derek’s what is this budget for supply chain? Is that a question valid? So there’s an authorization happening. And let’s say if all these are good, then the next thing that happens is it goes into a gateway layer. Now when you talk from a tokenomics, you’re just looking at a gateway, a model routing. But me as an architect, what did you. Go ahead.
[00:16:41] Speaker A: So, Rob, I think what you’re advocating is look at the big picture. Am I getting that right? Yes, yes, right. The tokenomics is just one piece of the pie. I think that’s going to be music to both Derek and Joanne’s ears because Derek’s going to talk to me about security and Joanne’s going to talk to me about data. And I do have a question for Joanna about something you said around harnesses and looping. But you could see in some of the numbers. I’m going to blow this, this research slide up. I’m not going to go through all the details, but it took me a while to digest the second bullet because I thought menlo Ventures had the best data on how much we’re spending on AI and they came back with a bottom up $37 billion number that AI is being spent on by enterprises.
Compare that to the $70 billion that OpenAI and Anthropic together have in their filings.
That of course is also including a lot of waste.
And that’s your point. Experiments go awry in keeping things running. It’s also including consumers. So that’s why we have a much bigger number going into AI and anthropics number. Now as you start getting into IDC’s number, you start looking at infrastructure. That’s your total cost there. And now it’s, you know, more than tenfold bigger. And then we start looking at global and you look at the entire end to end factors. That’s really Gartner’s almost $2.6 trillion being spent on AI. And so Derek, I mean like what Rahm is saying is we need to look at more than the 90% drop in token costs is everything that’s around it. That’s really driving the entire cost model that we need to consider when we’re building our AI out. Hi Derek.
[00:18:30] Speaker D: Good morning. Thank you. Yeah, and I agree. And Rom’s statements are right in line. Some of the things I was looking at also the measurement of these costs for the tokens.
I think businesses need to stop looking at measuring just the token cost, but look at the consumption behavior because that’s what’s really driving the cost. And if you’re looking at that, as Ram mentioned, the total cost of ownership, that’s a huge thing because now you’re going to stop measuring AI token costs and you look at what’s the outcome, what’s the cost per the customer that I’m serving, what’s the cost per report being generated, what’s the security investigation completed? All these things tie back into the cost per productivity gain. And I think when you look at all those things, the biggest picture, as Rom said, from an enterprise point of view, what business capabilities did AI improve and how much did it cost? So when you look at that, the things that come back to mind are, okay, well what things did you put in place? Have I created AI usage guardrails early? Have I made it so when people go to the these particular sites and make these queries, they’re not going to inject or put something of company propriety into those particular things? I’m looking at the models and the policies around the models. I’m looking at user quotas, I’m Looking at the prompt utilization and standards that have been put in place for people to now use AI securely and use it safely when they’re making these queries. I’m also looking at adrenal mention the monitoring. This is going to be huge because you need to know what your users are doing, where they’re going, how they’re using it. And it goes back to when you’re talking about resilience, kind of equivalent to a zero trust model. All these things need to be put in place. As I’m looking at establishing the AI values and standards, you know, what kind of scorecard am I keeping these go back to the metrics. How are you measuring expected savings, how you met, predicted or expected risk reduction? How are you measuring the expected revenue impact versus what you actually achieved? And then overall compliance.
You have the compliance standards you have with your regular operating usage in the company, business operations. But now you’re looking at are you still compliant? When you look at artificial intelligence usage and what are the implications around that? And it’s mentioned not just the, you know, the total life cycle cost, the total cost of ownership, all these things come into play when you’re trying to figure out is this going to be good when I’m using all these, these tokens or whatnot. And the things that Rob discussed, you know, these are spot on for things that. And the big picture a lot of companies don’t think about up, up front and by taking a step back and looking at what are you really looking at as far as business value and how am I going to gain that using these tokens, that’s what really needs to be a part of the objective moving forward.
[00:20:59] Speaker A: Thank you, Derek. Hi, Joanne.
Now you want to talk about return on data.
[00:21:05] Speaker E: I do want to talk about return on data.
You know, when you’re. I use a lot of manufacturing examples, but a lot of what I discuss in my posts about insight versus execution is applicable to any industry. Right. I use manufacturing examples because they’re very tangible, they’re very quantifiable and the impact is immediately absorbed by the reader. But overall, I agree with a lot of what has been said.
I would say two things. First of all, from a production point of view, the cost lives in the long tail. And the long tail is not only the harnesses and the loops. And I know you have a question about that.
But return on data is a metric that I created after about a year and a half of talking about time to data, time to decision, etc. Etc. It’s a framework and it’s a yield metric that is to provide anybody in the organization, whether it’s the CFO or the CEO at that level or even below it, a way to measure, you know, if you ask somebody how much or I spent $500 million on a piece of equipment, you can get an ROI on that based on a payback period and the usefulness, etc. Etc. The outcomes that are derived. But if you ask the same organization, what is the cost, what is the value created by your data estate and AI in particular, you get a lot of blank stares. And the reason for that is because there’s a percentage of data that goes into the cloud under the lift and shift paradigm that never gets analyzed. If you ask IBM about that in the Business Value Institute, they’ll tell you 90% never even gets analyzed. So you’re using roughly 10 to 22% of the data you’re storing in the cloud and, and you’re using even less for your AI.
And if you look at the cost at innovation in the lab, where things can go wrong and you’re in a sandbox environment, we’re not talking about anthropic and hugging face, but if you look at sandboxes, you can control the cost by how you do the prompt engineering, how you use cloud code. There are some tips and tricks. I put three of them in in the notes.
That being said, from the return point of view, you really need to look beyond the workflow, at the people, the process, the technology that’s being scraped. And two things of note. Own your AI, don’t outsource it to somebody else.
Look at small language models in a much different light. They’re much more cost efficient. They, they’re much more governable. You can do execution audits on business processes that you currently have, pinpoint where you’re deficient, and then build your AI capability in the small language models that allows you to really take advantage of the technology of AI without the extraordinary spend or the lack of control.
You need three things. You need agency, you need authority, you need accessibility.
Those three A’s, the accessibility being the security side of it, the authority being the accountability side of it. And the agency being is we as a corporation are allowing not only our employees, but the world to know that we have an AI capability. How that manifests is a variety of different outcomes. But at that point, return on data as a framework can basically say you’re currently using 8% of your data. You want to increase the utilization of your own internally sourced data which is protected by your security policies and give it a certain amount of agency within your four walls and build the trust needed to get to a point of autonomy. That’s another a that autonomy is earned. It has to be evidence based by repetition or source data. Run it on your erp, run it against your mes. If you’re in manufacturing or in financial systems versus what you’re looking at, do that comparison and do it with the capabilities that are currently available, whether it’s a weighted model, an unweighted model, meaning it’s open source, but keep it within your walls. You will get much more productivity, you will gain much more in the lack of margins leaking opportunity. Cost will go down and revenues will not erode the way they currently do now.
[00:26:05] Speaker A: Thank you, Joanne. No, no, there’s a lot there. I do have a couple follow up questions, but I want to bring in our other speakers first and I want to throw a prompt out so that I can get Heather and Liz commenting toward the back end. You know, when I think about AI and cost, number one, you know, I think measuring TCO is a fallacy. I think CIOs have wasted a ton of time and money trying to track down and become super accountants and capture everywhere where there’s cost. And I think for the most part it’s unrealistic and, you know, it’s just a diversion from actually getting things done.
And I’m, you know, throwing that softball over to Liz to comment on because I know she and I have been part of very similar exercises to get TCO that yielded nothing. And my prompt for Heather. Heather, my number one cost when it comes to AI is a people cost.
You know, I think we’re missing the big picture here that change management and training people and reskilling people is a real cost. I’ll let you comment on that when you raise your hand. John, your thoughts? We’re talking about innovation costs, production costs, and AI agent lifecycle costs.
[00:27:22] Speaker F: Welcome John Isaac, thank you for having me on. And the thing I was going to start with is if you really want to control cost, you have to manage it. And it really starts with providing the education on how you should use things. When should you use the most powerful models, when should you use the medium and the lightweight models?
And you have to really make it clear to people that maybe you want to use Fable or the sole models for their planning, and then you want to use maybe some of the other models for actually building things.
And that using the right model at the right time, you can save massive amounts of money.
And then the very Next thing is that you really can have runaway spend with this. And so if you don’t want to have runaway spend, you really have to set token limits or credit limits that are appropriate for kind of each role that people have.
People that are, that are not likely to be developing software that probably have a lot lower limits because they’re just going to be interfacing through the chat window. But where you can really spin the meters is when you’re doing software development and it’s easy to spend several hundred dollars in a couple hour session. I’ve done it myself.
And so I think you need to look at how are the people using it and set appropriate credit limits for things. And then I think you really have to say guidance on you can’t be using tons of credits if you’re just doing hobby projects. Like if you’re going to really use a ton of AI building something, it has to be for a real project.
And then you have to have cost governance software and you have to go actually back and look to see who’s using stuff and is it appropriate. And if you’re not doing those things, you’re going to get spent and who knows if it’s going to be adding value.
And you have to have cost governance across everything because you may have agents in production that are really spinning the meters.
And if you’re not looking at those two, you can’t just look at the users, you got to look at the agents too.
[00:29:15] Speaker A: Thank you, John. Let’s bring Joe in. Joe, you’re going to go for innovation for 200, production for 500 or AI agent lifecycle for 5,000?
[00:29:27] Speaker B: Well, as usual, I’m going to go for the fundamentals.
I don’t think anything we’ve talked about.
And by the way, I think Ram, Derek, Joanne and John have all brought out very, very interesting and valuable guidelines around how to manage this stuff. But managing it, in my view, requires transparency.
It requires, as John said, education.
You have to communicate to everyone up and down and laterally what these things are likely to cost, where they should be used, where they should not be used.
We often talk about celebrate the wins so people understand what true value creation is and learn from the failures where the cost outweighs the benefits.
Sharing information, communicating effectively is always key across the board.
[00:30:26] Speaker A: Joe, I have a question for you on this. It’s in my AI agent lifecycle and I want a ex CIO’s opinion on this before I lose it. And that is, you know, we have some folks who are just experimenting and some who have gotten agents into, into production. And now I’m hearing from CIOs, they’re getting into the hundreds of thousands, not hundreds of thousands, hundreds to thousands of AI agents in production.
Do we need to bring back chargeback models, departmental chargeback models, as we start thinking about the cost of running agents in production? The same way there are people chargeback models to departments for the number of employees. And Liz is giving me a heart that you cannot see.
I know where Liz is going to stand on this. So I want to hear Joe, what’s your opinion on this? Is that something CIOs have to consider?
[00:31:24] Speaker B: Absolutely. That is a very valid mechanism for containing sprawl or runaway use of any technology.
We certainly, we had this issue when people were rushing to the cloud. Oh my goodness. I can have, you know, really cheap unlimited storage. Somebody’s backing it up for me. It’s accessible from anywhere. Oh, lots of attractive features to cloud. And what happened? People put all kinds of data out in the cloud and then simply forgot about it. You know, so you, you do, you do have to put some sort of governance around it, whether you’re charging back the cost directly or using some other mechanism to contain that sort of runaway use of, of any technology. And that certainly applies to AI and agents in particular.
[00:32:14] Speaker A: Thank you folks. Welcome to this week’s Coffee with Digital Trailblazers. Our 181st episode. We’re talking today about Fin AI and managing AI costs where we’re looking at the different considerations in the innovation timeline, how we bring agents to production and what those cost factors are and then looking at what’s going to be happening when we get more and more agents into production and what those cost factors are. I want to thank Ram for being our enterprise Architect special guest. He’ll be speaking in a few minutes.
Folks, I left a link in the comment stream. I will put it in there. Again, I’m looking for your input around upcoming topics. I have a LinkedIn poll running around that.
Please do click on that sometime during this session today and share your input with me. It’ll take about five seconds. I had some food bars this morning with the research slides. I did not was not able to go through it in detail.
It is up on my Coffee with Digital Trailblazer page and that’s@drive.star cio.com coffee. You can find it there.
As well as being able to click into all the research that’s there, I want to start reminding you of upcoming topics. The 31st we’ll be talking about evaluating technologies in the AI era value risk, ecosystem evaluation. This is looking at how we make our platform decisions and what rules have changed with AI is a foundational capability, not an add. On the 7th, this is by listener demand, we’ll be talking about Networking in the AI Era, A guide for digital transformation leaders.
I’d love for all of you to share your questions about where you think you need help around networking. That will be our topic on the 7th and I will tell you that the 14th topic, I don’t have the exact title nailed down yet. We will be talking about building AI around broken and legacy processes and how we have to fix that. That will come out next week as well as all the august topics.
So please look out for that. Derek, your hand is up again. Are we talking production or agent life cycle? Where are you going?
[00:34:41] Speaker D: We’re talking production. We’re looking at what are some of the things that we learned through adopting this stuff Heavy quickly and it’s been kind of scary. You know we, we thought we learned when we migrated to the cloud and cloud automatically we reduced costs but we didn’t realize the, the cost for provisioning and the cost for security and the way that shadow AI exploded. You know the biggest thing I see right now with these production is people are implementing first and then putting governance on after the fact, which is a huge problem when you’re looking at the unused resources. As Joe mentioned earlier, the data sprawl, the cost overruns, it creates more security gaps and most of all you’re creating security risk. These things need to have more organizational discipline around it. When the deploying the copilot, deploying and building the agents, building the enterprise data. We need to establish these AI governance, these risk models, the data privacy policies and accountability as things that Joanne mentioned to. These are all important things that need to be done. As mentioned before, you can’t see it, you can’t manage it. You need visible accountability to drive the accountability of what’s going on on with your AI architecture. You know these things that now AI tools or employee use, AI tools and models consuming the budgets, AI tools being used to touch sensitive data. We need to figure out how can we see what’s actually taking place and not just we just rely on the older tools that we have. The tools that we have in place now need to be upgraded to support artificial intelligence applications, speeds and risk. We also look at the things regarding these tools based on the architecture. Everybody assumes you can just take artificial intelligence and throw it on the architecture you have. Yes, if it is secure, if it has the policies in place and things to govern it. The problem we’re seeing is now we don’t have the good architecture. We have excessive API calls, we have duplications of agent workflows, we have expanding attack surfaces and when I look at this now we really need to require from a finops point of view what are we looking at when it comes to fin AI? What are we looking at when it comes to AI security operations and AI security resilience as well as risk management? All these things are lessons learned that we need to look at sooner rather than later. But the problem is it’s still taking place. People are, some people are realizing the faults. They realize they need to go back and do it the right way. But other companies are still full throttle. The model that companies they did with the token buffet, that’s what I call it. Whereas a free for all that’s when they realized they did an oh moment or an aha moment. These are the kind of things we would need to adapt and think about earlier and not after the fact.
[00:37:07] Speaker A: Derek, I think you’re spot on with this. I mean we backed into governance and security on the cloud and sometimes that was just a lift. Very often it was a RE engineering exercise and all too often it was a RE architecture that was required with AI. I think your point is once the data is out there, once the leak is out there, it’s really very difficult if not impossible to back it out.
Once you’re in the headlines like OpenAI was around their the hack that happened this week and I’m not going to go through all those details but it’s a fascinating story.
Once it’s out there, it is out there and really how the back pedal through bad press and a real risk that materializes like that causing real damage. So I think that’s your point around that. Rom, welcome back. Are we talking production or agent life cycle? Where do you want to go right now? Rob?
[00:38:03] Speaker C: Both ways but Derek just put so many thoughts in me or so many dots like everything is so relevant and connected. The only few things that I’d say Derek and everybody is when we put production models or models into production the least that we are talking less is what does it mean from a cyber insurance. Right now cyber insurance seems like a vague, vague, vague way. There’s not a clear way of showing risk analysis or how do we build an ISO 27001-42001 into cyber insurance. And when a breach occurs cyber insurance guys are backing off and that’s a big, big concern. That I have and we need to bring this into forefront. People like us or Isaac, you should have more cyber insurance carriers talk about it right now. It’s so discretionary and so vague. My biggest heartbeat is do we understand at a boardroom level. So again, I know I’m transgressing a sec. If you have seen after cloud came there is this buzzword called cnapp. I’m sure many of you know this cloud native application platform. It has scaffolded into 14 sub platforms. We’re all giving the same story. I spoke to my former. How long can I keep telling the board, hey, it’s a risk avoidance. I don’t want to use the same caption. IBM says an average data breach is 4.4 million. But I’ve been using the same slogan again and again. I feel fed up. How do I show that by implementing all these.
And now agents have come. The new buzzword is adr. So to detect an agent response, you’re using adr. We are looking at the posture we’re calling aspm. But when I go to cyber insurance, say I’m putting all these, I’m sandboxing all. I don’t have any risk reduction in them. Not an audit trial. So I’m using risk reduction internally. But I’m not able to translate that into P and L other than saying that the cost, there’s a brand and this, that story I’ve been playing again and again for the last 10 years.
So my only request is all of us have to start speaking very loud and bold on cyber insurance and disaster recovery. For sure.
[00:40:09] Speaker A: Rahm, we’re going to have a topic come up on cyber insurance. No, it’s just a good prompt.
You know, it’s sort of like, it’s almost like that, you know, get out of jail free card from, from, from the old Monopoly games. Like I buy my cyber insurance and you know, I kind of walk away from all the risk.
And I think your point is no, you don’t particularly when it comes to AI. So, you know, if you have somebody to recommend from the carriers that could speak to this, we would need somebody with that background to host that conversation. I will take up that challenge and find someone.
Liz, I throw you a softball on chargebacks.
I, I want your comment on that.
But your overall we’re talking about cost here. You know, we know there needs to be a value equation.
Where do you want to go in the conversation, Liz?
[00:41:08] Speaker G: I just, I feel people over complicate total cost of ownership.
[00:41:13] Speaker A: Thank you.
[00:41:15] Speaker G: Over complicated.
I mean it’s just not that hard. You know that you’re going to do something with this thing that you’re purchasing and it’s going to have a bill, who’s going to pay it and whose budget is it going to. That’s it. And that person needs to be accountable and responsible for making sure that that is actually worth its while. That’s why God invented the pmo. And governance is not a four letter word. I’ve said it a million times.
It’s actually a support system to keep people from getting into trouble.
You know, if you actually have some rigor around what you’re investing, what your usage is and where that, where that usage is going, you can actually take that and use it as a way to charge back the, the dollars. It’s just not that hard.
[00:42:08] Speaker A: Well, I think if you keep life simple, Liz, I think you know, and draw the line on the things that start getting into micro accounting, I think you’re onto something.
I will tell you that when you start seeing platforms labeling AI governance capabilities, they are starting to enable per developer, per department, open level cost structures, model level cost structures and agent level cost structures. I got a demo of Calibra’s platform this week and saw that working.
So the platforms are starting to catch up with the ability to do this and you know, reason I’m Heather, I’m bringing this up because somewhere in here there’s a let’s if we just only look at the productivity angle and I know I want people to look beyond productivity. We only look at the productivity. We’re shifting from what people used to do to now. AIs used to do that. Cost is in people today we need to be able to factor in that change for people. But also what that’s been replaced with on the cost side with agents. I’m just wondering your thoughts around how we think about people in this.
[00:43:24] Speaker H: Well, I think people need to remain in the human resources as well as talent acquisition. They both relate to people.
There are certain things inherently that can be used in talent and acquisition departments that would rely on AI sourcing being one of them, validation being another. Doing the research before you even find candidates, all those things make sense. But when you start identifying the candidates and you need to speak with them, there’s absolutely nothing that can replace the human aspects. Even if you do a video chat in advance of an interview, I know a lot of people just turn off and they say that’s it, I’m going to withdraw from the whole process so it can be utilized across the board, especially at higher levels. So when you are going to keep some people and yes, you need them, you may even have to raise your price to get the rate that you’re paying the talent acquisition recruiters internally as well as externally because they’re doing more with less, they’re speaking to more people and the value that they bring is, is even more valuable.
The chargeback, any project that you do, as Liz mentioned, it’s not that difficult. If you’re hiring a PM project manager or a VP for a certain department, the chargeback should go back to that department.
There’s goals that are being set in talent acquisition teams and in HR in general that have to be met. So if you’re going to meet them, then charge them back. So that to me is simple. One of the votes that the areas that Isaac is requesting for a vote is on AI and hiring. So this would be a great way for us to chat about it. And just one last comment about disaster recovery. I was involved in that many years ago at JP Morgan. I think we did it twice a year, if that. Usually it was maybe even once a year. With the change that is going on in, in AI and in cybersecurity and so many other areas that impact a business. If you’re going to have disaster recovery, which I absolutely advocate, it’s got to be more frequently done more frequently and with greater value that is shared because changes are going to have to happen and you can’t wait until six or nine months later to make that change.
[00:45:53] Speaker A: The really good factors, you know we talked about apps going from a nuisance when they weren’t working to becoming mission critical and now AI having so many moving parts to it and now part of all of our operations and we’re not testing or even planning for Dr. Around it. It’s a very interesting data point Heather, that you bring up. Joanne, I want to hear your comments, but I have two questions I hope you can comment on. Number one, what do you mean by own? Your AI was a statement you’ve made a couple of times here. So I just want you to put some box around what that means and then I want you to comment on something Ramesh brought up around harnesses and looping. It seems to be there’s a cost and quality trade off here as we’re developing our agents, how many, how much looping, how much analysis we’re putting by default or within the bounds of what our agents can do and and if we underinvest that we’re going to have maybe poor results if we over engineer we’re going to have higher costs. I wonder if you can comment on how do we make that a little bit more data driven. Go ahead, Joanne.
[00:47:09] Speaker E: Okay, so first of all, some of it goes, and I don’t mean to sound disrespectful to anyone when I say this, some of it goes to prompt engineering.
And when you’re building agents, I mean, we have 11 layers in an agent. If I, if I were to stack it up like dot, you know, like boxes, there’s 11 layers. The thickest one is around context, which means it’s inference. And so when you design the agent, you really have to look at the harnesses, the environment around the agent. The looping is what happens on a failure. Right. And one of the standards that we push with every agent across the platform is if you don’t know, stop and defer to a human, that’s one way to lower your cost. But what it talks to also is the need for semantic layers, context layers, and divvying things up in a way that makes sense to the organization. For example, you may have an opinion, but if I don’t understand the approach, the intent or the perspective that you’re bringing, the agent called Isaac would fail on the loop.
I need to break things down in a very granular way to get success and avoid that trade off situation that you’re talking about or mitigate it down to 9010.
Because if I don’t do that every time that agent runs, it’s not being enhanced by your knowledge or experience, it’s not being enhanced by institutional knowledge. And so it’s going, it’s not learning. And because it’s not learning, it’s not improving. Therefore it’s also raising your cost, but not raising the value that it’s bringing to your organization. Does that make sense?
[00:49:19] Speaker A: Oh, it makes a ton of sense. I mean, I think this is, you know, one of those things that is, you know, one of these undertakings when you start building your own AI agents. And now you’re shifting from the simple, you know, use cases that you throw at it during innovation life cycles. Now you bring it into production life cycles and you haven’t planned for all that complexity.
And I think that’s an important lesson here because, you know, I compared this to moving to the cloud. You know, the cloud had a scaling factor, right? You tested in dev and test, you did some load testing, you scaled up, you put your scripting in place to auto ramp up and ramp down.
The dimensions were pretty finite. We’re talking here.
Inputs that are hard to manage.
And then you brought in a very simple rule. Bring in the human when required, put that right up front in the guardrails. You know that that’s a principle organizations have to learn.
[00:50:30] Speaker E: It is a principle that they have to learn. And the other thing about it is also that a good percentage of your cost on the tokenomic side is going to be around inference. And the more you can have your organization understand what that really means, which is multi perspectives. You know, think about a bunch of words that I related earlier.
Your intent, your, the content, the overall context, and then break it down into smaller chunks, no pun intended.
That’s the way to start designing your agents. And you’ll find that over the course of the development, moving from early innovation in a lab towards production, that context layer has many, many, many subsets. It’s like a, an 11 layer cake.
And if you don’t hit as many of them as you can, your loops are going to be wrong, your harnesses are not being effective, your costs will go up and the value that’s being created gets lost.
And so that’s where a lot of, a lot of things don’t move to production that you were kind of betting on would be in production.
[00:51:42] Speaker A: Thank you, Joanne, sorry for cutting you off there. I didn’t mean to.
We have a really, really good comment stream here today. Lots of ideas and I’ve gotten some of them into our D dashboard, but not all of them. Liz, you’re coming up next after John, you have this great quote up right on the screen. If we wait until we have quote, perfect governance models, we will not have any governance. Put something in place and refine.
Liz, you’re, I’m putting you on the spot. What are you putting in place? Just, let’s get John first. Give that some thought. John?
[00:52:13] Speaker F: Well, I, I want to just say one, one comment is a lot of times people, people talk about the cloud cost and saying that like once people started using cloud, the cost really exploded. But, but when we were using data centers, we had no idea what the costs were. And so once we actually started moving things to clouds, we could actually see what the costs are.
And then now it’s the same way that now that we’re starting to use AI in the beginning, when we’re using AI through web browsers, the costs are by and large pretty trivial. And so to Joanne’s point, if you have your own AI, all the costs are on building the models. The inference is a lot cheaper usually compared to actually building the model. But then if you’re, if you’re using commercial models, the costs that people have that are really exploding are almost all on the development side. And so it’s just like if there’s any place to really put the governance, really put the education is how do people very carefully do software development using AI? Because that’s, that’s the thing that’s going to spin the models more than anything. If you’re using it in production, you, you can at least that’s, that’s a lot more, that’s the, the usage on the production will scale as people are using your websites or your software. And so that’s going to hit you over time. But if you’re, if you want to looking for the spikes, it’s, it’s all on the software development side.
[00:53:35] Speaker A: All right, John, I’m throwing that softball. Liz, I love this comment, but then
[00:53:41] Speaker G: when you give me a softball, then I’m like, well then just try something.
[00:53:47] Speaker A: I’m going to give you my answer. My answer is who is responsible for this?
[00:53:50] Speaker G: I have more to say, but go ahead.
[00:53:53] Speaker A: Who is responsible for this AI agent’s decisions? If you can’t get that answered before you start contemplating it, don’t start. That’s my first governance role. Go ahead.
[00:54:02] Speaker G: Yeah, so listen, I, I’m in charge of governance in my current role and we do a very lightweight cost benefit analysis for any initiative that comes through. And the rule is if it’s going to impact resources or it’s going to impact dollars, then it has to have some kind of cost benefit so that even if we know it’s going to lose money, it’s a conscious governance decision that we’re going to do it anyway or not. Right. So the key here is not to over complicate dollars.
Estimate something, try anything. Right? You think that they’re going to explode at this rate. Put out how long you think that that’s going to be useful for, for how many years, five years, whatever. Throw something on the paper and then measure it. The key thing is to actually circle back six months after production, maybe in three months after the production and see what you’re really getting out of it. So the government’s model that, that I put in place where I work, it has not only this very lightweight cost benefit analysis upfront, it’s a five year payback model, super light. And then on the back end, once the project is complete and we have a closeout and a go light, everybody’s so happy, blah, blah, blah. I have put on the calendar a six month look back and that six month look back has in it a specific metric that was promised when it was approved that we’re going to go and take a look at to see if we actually hit that mark and it gives us an opportunity to adjust. Either it’s a crappy project and maybe we should have never gone there or we need to fix it so we actually do get there. Our model was bad and we need to fix it. Going forward, any of those things could be true.
[00:55:48] Speaker A: Thank you, Liz. I’m just going to go around the horn because we got four minutes left. I want to hear from Joanne and give Rahm some last parting thoughts. Go ahead Joanne.
[00:56:00] Speaker E: I just wanted to make a quick comment on something that you posted on the AI agent life cycle chargebacks. There’s a quote from Rick Allaire.
Governance is an afterthought in AI. It was similar in the data world, but risks are much higher. If governance is a bolt on or an afterthought, don’t bother because governance has to be part of your design process and it has to be initiated at business process, either redesign or workflow redesign. If you don’t do it, then you will not succeed with AI. And governance, by the way, is more than just the cost. Governance is also making sure that agents follow policy decisions, that there’s industry standards that are applied. They have to meet the conditions of both of those. If they don’t and it’s not part of your AI or agentic AI initiative, then you’re going to fail.
[00:57:05] Speaker A: John, it sounds like we have a lot of work for digital trailblazers out there, you know, who want to make sure their organizations are experimenting and moving AI agents into production wisely.
There’s just a lot of concern here that we’re bringing up. Rick has this great quote here. Uber didn’t learn its lesson from from the cloud.
They had sticker shock. They did it again with token consumption. Just, you know, just this week, a lot of stories about the lack of governance becoming a problem.
I’m going to go to John very quickly and then ROM to close up.
[00:57:42] Speaker F: Yeah, and if you look at the P and L and the statements on the AI companies, we’re really not paying the cost of AI and it’s just like Uber. When Uber rolled out, I used to be able to take a Uber from Seattle, Bellevue at almost, almost nothing and they were really just blowing through a bunch of investment money. So now we’re playing the true cost of Uber and in the future we’re going to pay the true cost of AI. It’s going to be a lot more than it is now.
[00:58:05] Speaker A: John, I was sticker shocked when I saw like this range of 30 billion to multi trillion dollars of spend on AI and then anthropic and OpenAI sitting on this sort of minuscule number in the end of $70 billion of revenue right now.
Wow, this isn’t going to add up. There’s, you know, somewhere in here enterprises, the market, the, you know, the frontier models are just going to fall off the cliff. It’s just, it’s just wild. I want to give Ram the, the closing thoughts. Ram, thank you for joining us today.
[00:58:43] Speaker C: Yeah, sure. Two, three things. Great, great topics.
You know, when I hear change management, we say change management for people. I have a humble request. Change management has to be for the leaders and I’m hoping people means the stakeholders. Just because you’re the CEO doesn’t mean you know everything.
This is the first time change management should include every stakeholder. It’s been missing again and again. My team needs this. My need. I have a vision. No, even you need it. When a contingency occurs, it’s just not a chief exit, it’s the CEO because the brand is at stake.
The second thing that I want to say is if Uber CEO has made a statement that, gosh, are we really getting the value for AI?
If this is a statement coming from the CEO, watch it out. And the third thing I want to say is Satya Nadala has made a statement saying that humans are cheaper than AI. So if we have said AI, if we have said tco, why are these executives coming out? Which is by the way a great thing they’re acknowledging is exactly what McKinsey is saying. Only 5% is making P and L. We don’t want to complicate, but if you don’t understand the depth to what Joanna said, The 11 layers, it’s an engineering task. But if you don’t think of the 11 layers and you just think input to output, you can be happy now, you can celebrate it. But you got to put in back of your mind these 11 layers matter a lot.
So if you’re designing architecture for scale, design for future, not just for today and celebrate it, celebrate success, but celebrate for a long time. But be open based.
[01:00:16] Speaker A: I love these takeaways. I mean we had a brainstorming the other advisors in me a little over a week ago. We talked about leaving everybody here with some concrete action items and you just kind of nailed it.
So thank you for that. Thank you for joining this week.
Thanks to all the Advisors at the Coffee with Digital Trailblazers. Our topic next week of evaluating technologies in the AI Era, Value, Risk and the Ecosystem will cover it on those three dimensions.
The seventh we’ll be talking about Networking in the AI Era Guide for Digital Transformation Leaders and then I’ll have the remaining topics out this week.
Sign up for my newsletter. It’s starcio.com/driving-digital that’s where you’ll first see it announced. You’ll also see it on the Coffee with Digital Trailblazer page that drive.starcio.com/coffee and then if you have not voted for your topic, visit the URL in the comment Stream. It’s a LinkedIn.com feed URL. I can’t read it off, it’s just too long. But do find it in the comments stream. Click on it. It’ll take you two seconds to pick one of four Folks. Great conversation today. We’ll be back next week talking about evaluating technologies in the AI era. Everybody have a great weekend.

























Leave a Reply