Spend just one day in New York City, and you might be transferring subway trains, hailing taxis, calling Ubers, sitting in traffic, or scrolling the streets on Google Maps for hours. Each one of these interactions constitutes a decision and data point, produced by just one person. Multiply that by a population of 8 million more, on top of tourists, and the busiest city in the U.S. generates millions upon millions of decisions and data points every single day.

    New artificial intelligence models that parse through that vast sea of human-generated data, however, could further expedite and smooth the decision-making process, from commute times to traffic safety. This is the crux of Assistant Professor Yingxue Zhang’s recent National Science Foundation CAREER Award, for which she won a five-year grant of $584,649. 

    “In urban life, people make decisions every day. Taxi drivers want to pick up passengers, and if people want to go to work, they need to switch to different public transits,” said Zhang, a faculty member at the Thomas J. Watson College of Engineering and Applied Science’s School of Computing. “They make decisions every day, every hour, every second. We want to model this kind of process to make the whole decision-making process more efficient and effective.” 

    The CAREER Award is the NSF’s most prestigious distinction, given to early-career faculty who are poised to become the future leaders in research and education in their respective fields. Zhang’s project will be the first to use a particular type of deep learning technique, called offline reinforcement learning, to wrangle the ever-changing dynamics of urban life and spaces.

    “This technique has never been applied in the urban domain, which is a spatial-temporal domain. The data is spatial-temporal data. We have spatial correlations, and we have temporal correlations. It’s much more complicated than the ordinary robotics domain,” Zhang said.

    Offline reinforcement learning is a branch of the more basic “reinforcement learning.” Zhang compares reinforcement learning itself to the way a robotic vacuum maps out the layout of a room as it begins to clean. The Roomba observes the environment around it to make decisions about how to navigate around furniture, walls, or obstacles — a modeling process which forms what researchers call “policy.” 

    But the vacuum doesn’t start out knowing everything about its surroundings. It has to learn its environment, and the way it learns is similar to the way we all do, through trial and error. It explores the room and bumps into corners and makes many mistakes — and in the process, it learns enough that next time, those things won’t happen again.

    Offline reinforcement learning, the technique that Zhang aims to use, is based directly on data rather than experience, since learning from trial and error could be dangerous in urban life.

    “We cannot tolerate a model making a lot of mistakes in urban life,” Zhang said. “So, of course, reinforcement learning is not a good option if we apply it in the urban decision-making process. That’s why we use offline reinforcement learning.” 

    The data in question comes from, essentially, everyday life: GPS data, traffic data, public transportation data. The catch, however, is that humans are messy, and the data we generate will also inevitably be a mess.

    “Because we use real-world data generated by humans,  GPS, and vehicles, sometimes the data quality is a problem,” Zhang said. “Sometimes the data is heterogeneous. It’s generated from different sources, different people. If the data is very diverse, how can we reliably model the decision-making process?”

    Though Zhang is working with data in bulk, each individual person generates only a little bit. Plus, data is not generated in a vacuum. The millions of human-to-human interactions that take place in any given city will also fuzz the decision-making process. Part of her CAREER project will be figuring out a way to generate a powerful model that can work off a lot of very little. 

    Moreover, she anticipates that whatever models arise will face a translation challenge: the environment where researchers collect their data is not necessarily the same as the real and changing thing. This “distribution shift,” she said, means the model isn’t directly applicable to the real world without some sort of intervention.

    “If we can integrate offline reinforcement learning into the spatial-temporal urban domain, then that would be great. It can help us solve a lot of practical problems,” she said. “But this kind of integration is not very simple, because we need to target a lot of unique challenges.”

    Luckily, Zhang has some safeguards in place, including plans to test the model in simulated environments. She will be collaborating with partners at the University of Maryland College Park and University of Pittsburgh as well as in Hong Kong to trial and incorporate these policies.

    Combining natural human language with the model through machine learning, she added, is also another way to improve reliability and task completion. 

    “For example, if we have a robot arm, we can provide a description: There is a robot arm, and the goal of this arm is to pick up a tomato. You can describe where the tomato is and its environment, so you can provide more details and descriptions about this environment and task and object,” she said. 

    Making the greatest impact

    Zhang traces her interest in urban environments and smart cities to her doctoral work at Worcester Polytechnic Institute. At the time, she had the choice between pursuing more conventional computer science routes or the lesser-explored area of spatial-temporal urban research. 

    “At that time, I was more interested in new techniques, like deep learning and AI. At that time, AI was at a very early stage, but deep learning had matured,” she said. “I was more interested in applying deep learning techniques to solve practical problems. That’s why I chose this direction.”

    It was Binghamton University’s status as an R1 ranked public university, as well as the caliber of its students, which drew Zhang to join faculty at the School of Computing. Its growing status as a leader in AI research also came as a serendipitous advantage for Zhang’s own work. 

    “Because we are working in AI, computational resources are very important. We need GPUs, and now we have a lot of HPCs,” she said. “We have high performance computing resources, so from my perspective, I feel like everything is ready here.”

    The results of her project will not only help build smarter cities, but also lay the foundations for new curriculum and courses at Binghamton. Zhang hopes to develop a new course related to offline reinforcement learning and its practical applications in urban life. 

    Through outreach to K-12 students and educators as well as potential collaborations with industry partners, Zhang also hopes to bolster interest and future workforce development in the field of AI and machine learning. 

    “We want to attract more people or students to get involved in this area, so that’s why we think workforce development is very important. We hope we can get more students — graduates, undergraduates, PhD students — to work in this area. And we hope that we can arouse the interest of K-12 students in this area,” she said. “In the future, if we can collaborate with industry partners, we can hold workshops as well. We hope we can benefit this area from different perspectives.”

    Her hope, she said, is that any resulting model will be usable for everyone. This model, built from ordinary people’s data, will be accessible to researchers and those same city-going humans, every day. For this reason, all results from the research, by the end of its five-year grant, will be open source.

    “Everybody can get access to the model and data, so they can replicate it. They could enhance the model and improve its performance,” Zhang said. “This is a good thing. We just want to benefit the whole community to the largest extent possible.”

    Share.

    Comments are closed.