Saturday, December 7, 2013

Not just micro displays: adding flat panel products to the mix

After about a decade of making virtual reality goggles and other near-eye devices based on micro displays, my company started demonstrating and shipping goggles that are based on flat-panel displays such as those found in smartphones. Many that have visited the I/ITSEC training show in Orlando had a chance to experience some of our new offering, and we were very happy with the feedback received.

It was not a difficult decision. Given the increasing resolution, diverse supplier base and lower cost of flat panel displays (as opposed to OLED micro displays), it made sense to start applying our innovation and our expertise in goggles to this new display technology.

To date, we have exclusively used OLEDs from eMagin. There are many good folks at eMagin, and they have nice products, but given their well-documented delivery challenges and new designs made possible by flat-panel displays, they won't be our exclusive display supplier anymore.

Where and when is it best to use OLED micro displays?

  • Where physical space is limited, such as when building a simulated rifle scope
  • When the contrast and response time of OLEDs are a must (that is, until OLED flat panels become widely available for goggle use)
  • Where harsh environmental conditions - especially temperature - are required, such as in our ruggedized HMD for training
  • Where high pixel density is required
  • Where being able to purchase replacement parts for many years is important
  • When low power consumption is critical

Where and when is it best to use flat panel displays?
  • Where wide field of view is particularly important
  • When it is important to have supplier diversity
  • When cost is a major factor
Not married to one display technology or another, it is now possible to choose the best one for each new product.


Monday, November 11, 2013

Do you make your own [...]?

Quite often, I get asked whether my company makes specific components - motion trackers, electronics, optics, etc. - that we use inside our virtual reality goggles.

As one would expect, we look at these 'make vs. buy' decisions individually, and ask several questions:

  • Can we add value to our customers if we 'make'? Can we generate a product that is significantly better or lower-cost or offers some other unique benefits relative to the 'buy' alternative?
  • How many of these do we expect to make? We'd be much more likely to buy when only small quantities are required and more inclined to make when there are more units.
  • Can we afford it?
  • Can we build it on time?
  • Can we create value for our shareholders by generating valuable patent filings or know-how?
  • Is this a discipline that we need to understand very well for our future business?
  • Does a 'buy' option exist?
Historically, these have been our answers:
  • Orientation trackers: we typically buy, but we then try to improve what we buy. We have worked with many of the leading orientation tracking vendors - Intersense, Intertial labs, Hillcrest, YEI - and have decided against developing our own. However, we have often worked with these manufacturers to introduce new features in their products or to optimize them for HMD use. We have added some of our own features such as predictive tracking when those did not exist. Last, we prefer to encapsulate the vendor-specific API with a standard Sensics interface because it allows our customers the benefit of maintaining their software investments when we change the motion tracker vendor inside our products.
  • Optics. To date, we have always designed and made optics ourselves. We have made optics for small and large displays, using glass, plastic and fiber, using a variety of manufacturing technologies. We believe that the our portfolio of optical designs is an advantage, and that optics are a critical part of the goggle experience.
  • Electronics. We often design our own electronics. Sometimes, we need special high-speed processing, and in other instances, we feel that we need something beyond simple driving of a display. This can be unique video processing, distortion correction or packaging that supports a particularly compact design.
  • Displays. We buy. We don't have the know-how nor the capital to make our own displays and in the world of changing display technologies, we're glad not to be locked into a specific one. Having said that, we have worked with eMagin in prior years to modify the size of one of their OLED driver boards to make a system more compact and achieve better optical design. It was a financial investment, but we felt we added value to our customers.
  • Mechanical design. We rarely design accessories such as helmet-mounts, but we do love to design goggle enclosures whether to give it our unique 'look', or to include innovative features such as hand tracking sensors.
  • Software. We write our own (or pay to have it written). Our software is so deeply tied to the unique functionality of our designs that it is not available for off-the-shelf purchase.
If you are a manufacturer and would like to see how we can use some of our technologies to help you get new, innovative products to market on short order, drop me a line.


Sunday, November 3, 2013

An Interview with Sebastien Kuntz, CEO of "I'm in VR"

Following my blog post "Where are the VR abstraction layers" I had an opportunity to speak with Sebastien Kuntz, CEO of "I'm in VR" a Paris-based company that is attempting to create such layers. I've known Sebastien for several years, since he was part of Virtools (now Dassaut Systemes) and it was good to catch up and get his up-to-date perspective.
Sebastien Kuntz, CEO of "I'm in VR"

Sebastien, thank you for speaking with me. For those that do not know you, please tell us who you are and what you do?
My name is Sebastien Kuntz and I am CEO of "i’m in VR". Our goal is to accelerate the democratization of virtual reality and towards that goal we created the MiddleVR middleware software product. I have about 12 years of experience virtual reality, starting at the French Railroad company working on immersive training, and continuing as the lead VR engineer in Virtools, which made 3D software. After Virtools was acquired by Dassault Systemes, I decided to start my own company and that is how "i’m in VR" was born.
What is the problem that you are trying to solve?
Creating VR applications is complex because you have to take care of a lot of things – tracking devices, stereoscopy, multiple computers synchronization (if you are working with a Cave), interactions with the virtual environment (VE).
This is even more complex when you want to deploy the same application on multiple VR systems - different HMDs, VR-Walls, Caves ... We want developers to focus on making great applications instead of working on low-level issues that are already solved.
MiddleVR helps you in two ways:
  • It simplifies the creation of your VR applications with our integration in Unity. It manages the devices and cameras for you, and offers high-level interactions such as navigations, selection and manipulation of objects. Soon we will add easy-to-use immersive menus, more interactions and haptics feedback.
  • It simplifies the deployement of your VR applications on multiple VR systems: the MiddleVR configuration tool helps you easily create a description of any VR system, from low-cost to high-end. You can then run your application and it will be dynamically reconfigured to work with your VR system without modification.
MiddleVR is an easy-to-use, modern and commercial equivalent of Cavelib and VR-Juggler.
How do you provide this portability of a VR application to different systems ?
MiddleVR provides several layers of abstraction.
  • Device drivers: the developers don't have a direct access to native drivers, they have access to what we call "virtual devices", or proxy devices. The native drivers write tracker data like position and orientation directly in such a "virtual device". This means that we can change the native drivers at runtime while the application is still referencing the same "virtual device".  
  • Display: all the cameras and viewports are created at runtime depending on the current configuration. This means your application is not dependent on a particular VR display.
  • 3D nodes: Most of the time the developer does not care about the information from a tracker, he is more interested in the position of the user's head or hand for example. MiddleVR provides a configurable representation of the user, whose parts can be manipulated by tracking devices. For example the Oculus Rift orientation tracker can rotate the 3D node representing the user's head, while a Razer Hydra can move the user's hands. Then in your application you can simply ask "Where is the user's head ? Is his hand close to this object ?", which does not rely on any particular device. This also has the big advantage of putting the user back in the center of the application development!
  • Interactions: At an even higher level, the choice of an interaction technique is highly dependent on the characteristics of the hardware. If you have a treadmill you will not navigate in a VE in the same way as if you only have a mouse, or a joystick, or if you want to use gestures... The choice of a navigation technique should be made at runtime based on the available hardware. In the same way, selecting and manipulating an object in a VE can be made very efficient if you use the right interaction techniques for your particular hardware. This is work in progress, but we would like to provide this kind of interactions abstraction. We are also working on offering immersive menus and GUIs based on HTML5.
Will this interaction layer also allow you to define gestures?
Yes, the interaction layer will certainly allow you to define and analyze gestures. Though this particular functionality is not yet implemented in the product, you will be able to integrate your own gestural interactions.
Do you extend Unity's capabilities specifically for VR systems ?
Yes we provide active stereoscopy, which Unity cannot do.
We also provide application synchronization on multiple computers, which is required for VR systems such as Caves. We synchronize the state of all input devices, Unity physics nodes, the display of new images (swap-buffer locking) and left/right eye images display in active stereo(genlocking). As mentioned, we will also offer our own way of creating immersive menus and GUIs because Unity’s GUI system has a hard time dealing with stereoscopy and multiple viewports. 
Do you support other engines other than Unity?
MiddleVR has been created to be generic, so technically it was designed to be integrated into multiple 3D engines, but have not done so yet. It’s the usual balance of time and resources. We made several promising prototypes though.
How far are we from true ‘plug and play’ with regards to displays, trackers and other peripherals?
We are not there yet. You need a lot more information to completely describe a VR system than most people think.
First, there is no standard way to detect exactly which goggle, tv, projector or tracker is plugged in. [Editor's note: EDID does provide some of that information]
Then in a display (HMD, 3D monitor, CAVE), it is not enough to describe resolution and field of view. You also need to understand the field of regard. With an HMD you can look in all directions. With a 3D monitor or most CAVEs, you are not going to be able to see an image if you look towards the back. The VR middleware needs to be aware of this and allow interaction methods that adapt to the field of regard. Moreover you have to know the physical size of the screens to compute the correct perspective. 
I believe we should describe a VR system not based on its technical characteristics such as display resolution or number of cameras for optical tracking, but rather in terms of what those characteristics means for the user in terms of immersion and interaction! For example:
  • What is the end-to-end latency of the VR system for each application? This will directly influence the overall perception of the VE.
  • What is the tracking volume and resolution in terms of position and orientation? This will directly influence how the user interacts with the VE: we will not interact the same with a Leap Motion which has a small tracking volume, or with the Oculus Rift tracker which can only report orientations or with 20 Vicon cameras able to track a whole room with both positions and orientations.
  • What is the angular resolution of the display? If you can't read a text from a given distance, you will have to be able to get the text closer to you. If you can read the text because your VR system has a better angular resolution, you don't necessarily need this particular interaction.
  • What is the field of regard ? As discussed above this also influences your possible actions.
The user's experience is based on its perceptions and actions, so we should only be talking about what is influencing those parameters. This requires a bit more work because they are highly dependent on the final integration of the VR system.
We are not aware of standards work done to create these ‘descriptors’ but we would certainly support such effort as it would benefit the industry and our customers.
Are most of your customers today what we would call ‘professional applications’ or are you seeing game companies interested in this as well?

Gaming in VR is certainly gaining momentum and we are very interested in working with game developers on creating this multi-device capability. We are already working with some early adopters.
We are working hard to follow this path. For instance, we are about to release a free edition of MiddleVR based on low-end VR devices and would like to provide a new commercial licence for this kind of developments. This is in our DNA, this is why we were born! We want to help the current VR democratisation. 
When you think about porting a game to VR, there are two steps (as you have mentioned in your blog): the first one is to adapt the application to 3D, to motion tracking ,etc. This is something you need to do regardless of the device you want to use.
The second is to adapt it to a particular device or set of devices. We can help with both, and especially with the 2nd step. There will be many more goggles coming to market in the next few months. Why just write for one particular VR system when you can write a game that will support all of them ?
What’s a good way to learn more about what MiddleVR can do for application developers? Any white paper or video that you recommend?
Our website has a lot of information. You can find a 5 minutes tutorial, and here another video demonstrating the capabilities of the Free edition.
Sebastien, thank you very much for speaking with me. I look forward to seeing more of I'm in VR in action.
Thank you, Yuval

Monday, October 21, 2013

Can the GPU compensate for all Optical Aberrations?

Photo Credit: <a href="http://www.flickr.com/photos/55514420@N00/5192375946/">davidyuweb</a> via <a href="http://compfight.com">Compfight</a> <a href="http://creativecommons.org/licenses/by-nc-nd/2.0/">cc</a>
Photo Credit: davidyuweb via Compfight cc
As faster, newer Graphics Processing Units (GPUs) become available, graphics cards can perform real-time image transformations that were previously relegated to custom-designed hardware. Can these GPUs overcome all the important optical aberrations, thus allowing HMD vendors to use simple, low-cost optics?

The short answer is: GPUs can overcome some, but not all aberrations. Let's look deeper into this question.

Optical aberrations are the result of imperfect optical systems. Every optical system is imperfect, though of course some imperfections are more noticeable than others. There are several key types of aberrations in HMD optics which take an image from a screen and pass it through viewing optics:
  • Geometrical distortion, which we covered in a previous post would cause a square image to appear curved. The most common variants are pincushion distortion and barrel distortion.
  • Color aberration. Optical systems impact different colors in different ways, as can be seen in a rainbow or when light passes through a prism. This results in color breakup where a white dot in the original screen breaks up into its primary colors when passing through the optical system.
  • Spot size (also referred to as astigmatism), which shows how a tiny dot on the original screen appears through the optical system. Beyond the theoretical limits (diffraction limit), imperfect optical systems cause this tiny dot to appear as a blurred circle or ellipse. In essence, the optical system is unable to perfectly focus each point from the source screen. When the spot size becomes large enough, it blurs the distinction between adjacent pixels and can make viewing the image increasingly difficult.
The diagram below shows an example of the spot size and color separation on various points in the field of view of a certain HMD optical system. This is shown for the three primary colors, with their wavelengths specified in the upper right corner. As you can see, the spot size is much larger for some areas than others, and colors start to appear separated.


Which of these issues can be corrected by a GPU, assuming no practical limits on processing power?

Geometrical distortion can be corrected in most cases. One approach is for the GPU to remap the image generated by the software so that it compensates for known optical distortion. For instance, if the image through the optical system appears as if the corners of a square are pulled inwards, the GPU would morph that part of the image by pushing these corners outwards. Another approach is to render the image up-front with the understanding of the distortion, such as the algorithm covered in this article about an Intel researcher.

Color aberration may also be addressed, though it is more complex. Theoretically, the GPU can understand not only the generic distortion function for a given optical system, but the color-specific one as well, and remap the color components in the pixels accordingly. This requires understanding not only the optical system but also the primary colors that are being used in a particular display. Not all "greens", for instance, are identical

Where the GPU fails is in correcting astigmatism. If the optical system causes some parts of the image to be defocused, the GPU cannot generate an image that will 're-focus' the system. In simpler optics, this phenomena is particularly noticeable away from the center of the image. 

One might say that some defocus in the edge of an image is not an issue since the central vision of a person is much better then the peripheral vision, but this argument does not take into account the rotation of the eye and the desire to see details away from the center.

Another discussion is the cost-effectiveness of improving optics, or the "how good is good-enough" debate. Better optics often cost more, perhaps weigh more, and not everyone needs this improved performance or is willing to pay for it. Obviously, less distortion is better to more distortion, but at what price?

Higher-performance GPUs might cost more, or might require more power. This might prove to be important in portable systems such as smartphones or goggles with on-board processors (such as the SmartGoggles), so fixing imperfections on the GPU is not as 'free' as it might appear at first glance.

HMD design is a study in tradeoffs. Modern GPUs are able to help overcome some imperfections in low-cost optical systems, but they are not the solution to all the important issues.


For additional VR tutorials on this blog, click here
Expert interviews and tutorials can also be found on the Sensics Insight page here

Monday, October 14, 2013

Where are the VR Abstraction Layers?

"A printer waiting for a driver"
Once upon a time, application software included printer driver. If you wanted to use Wordperfect, or Lotus 1-2-3, you had to have a driver for your printer included in that program. Then, operating systems such as Windows or Mac OS came along and included printer drivers that could be used by any application. Amongst many other things, these operating systems provided abstraction layers - as an application developer, you no longer had to know exactly what printer you are printing to because the OS had a generic descriptor that told you about the printer capabilities and provided a standard interface to print.

The same is true for game controllers. The USB HID (Human Interface Device) descriptor tells you how many controls are in a game controller, and what it can do, so when you write a game, you don't have to worry about specific types of controllers. Similarly, if you make game controllers and conform to the HID specifications, existing applications are ready for you because of this abstraction layer.

Where are the abstraction layers for virtual reality? There are many types of VR goggles, but surely they can be characterized by a reasonably simple descriptor that might contain:

  • Horizontal and vertical field of view
  • Number of video inputs: one or two
  • Supported video modes (e.g. side by side, two inputs, etc.)
  • Recommended resolution
  • Audio and microphone capabilities
  • Optical distortion function
  • See through or immresive configuration
  • etc
Similarly, motion trackers can be described using:
  • Refresh rate (e.g. 200 Hz)
  • Capabilities: yaw, pitch, roll, linear acceleration
  • Ability to detect magnetic north
  • etc,
Today, when an application developer wants to make their application compatible with a head-mounted display, they have to understand the specific parameters of these devices. The process of enhancing the application involves two parts:
  1. Generic: change the application so that it supports head tracking; add two view frustums to support 3D; modify the camera point; understand the role of the eye separation; move the GUI elements to a position that can be easily seen on the screen; etc.
  2. HMD-specific: understand the specific intricacies of an HMD and make the application compatible with it

If these abstraction layers widely existed, the 2nd step would be replaced by supporting the generic HMD driver or head tracker driver. Once done, the manufacturers would need to write a good driver and viola! users can start using their gear immediately.

VR application frameworks like Vizard from WorldViz provide an abstraction layer, but they are not as powerful as modern game engines. There are some early efforts such as I'm in VR to provide middleware, but I think a standard for an abstraction layer has yet to be created and gain serious steam. What's holding the industry back?

UPDATE: Eric Hodgson of the Redirected Walking fame reminded me of VRPN as an abstraction layer for motion trackers, designed to provide a motion tracking API to applications either locally or over a network. As Eric notes, VRPN does not apply to display devices but does abstract the tracking information. I think that because of it being available on numerous operating systems, VRPN does not provide particularly good plug-and-play capabilities. Also, it's socket-based connectivity is excellent for tracking devices that, at most, provide several hundred lightweight messages a second. To be extended into HMDs, several things would need to happen, including:

  • Create a descriptor message for HMD capabilities
  • Plug and play (which would also be great for the motion tracking)
  • The information about HMDs can be transferred over a socket, but if the abstraction layer does anything that is graphics related (in the same way OpenGL or DirectX abstract the graphics card), it would need to move away from running over sockets.


Sunday, October 6, 2013

Is Wider Field of View always Better?

I have always been a proponent of wide field of view products. The xSight and piSight products were revolutionary when they were introduced, offering a combination of wide field of view and high resolution. There is widespread agreement that wide field of view goggles provide greater immersion, and allow users to perform many tasks faster and better.
Johnson's Criteria for
detection, recognition and identification -
from Axis Communications

But for a given display resolution, is wider field of view always better? The answer is 'No' and thinking about this question provides an opportunity to understand the different set if requirements between professional-market applications of virtual reality goggles (e.g. military training) and gaming goggles.

Aside from the obvious physical attributes - pro goggles often have to be rugged - the professional market cares very much about pixel density (or the equivalent pixel pitch) because it determines the size and distance of simulated objects that can be detected. For instance, if you are being trained to land a UAV, or trying to detect a vehicle in the distance, you want to detect, recognize and identify the target as early as possible and thus as far away as possible. The farther the target appears away, the fewer pixels it occupies on the screen for a given pixel density.

The question of how exactly many pixels are required was answered more than 50 years ago by John B. Johnson in what became known as the Johnson Criteria. Johnson looked at three key goals:
  • Detection - identifying that an object is present.
  • Recognition - recognizing the type of object, e.g. car vs. tank or person vs. horse.
  • Identification - such as determining the type of car or whether a person is a male or a female.
Based on extensive perceptual research, Johnson determined that to have a 50% probability that an observer would discriminate an object to the desired level, that object needs to occupy 2 horizontal pixels for detection, 8 horizontal pixels for recognition and 13 horizontal pixels for identification.

Let's walk through a numerical example to see how this works. The average man in the United States is 1.78m tall (5' 10") and has a shoulder width of about 46cm (18"). Let's assume that a simulator shows this person at a distance of 1000 meters. We want to be able to detect this person inside an HMD that has 1920 pixels across.

46 cm makes an angle of 0.026 degrees (calculated using arctan 0.46/1000). At a minimum, we need this angle to be equivalent to two pixels. Thus, the entire horizontal field of view of this high-resolution HMD can be no more than 25.3 degrees for us to achieve detection. If the horizontal field of view is more than that, target detection will not be possible at these simulated distances.

Similarly, if we wanted to be able to identify that person at 100 meters, these 46 cm would make an angle of 0.26 degrees so the horizontal field of view of our high-resolution 1920 pixel HMD can be no more than 38.9 degrees. If the horizontal field of view is more than that, target identification will not be possible at these simulated distances.

Thus, while we all love wide field of view, thought must be put into the field of view and resolution selection depending on the desired use of the goggles.

Notes:

  • Johnson's article was "John Johnson, “Analysis of image forming systems,” in Image Intensifier Symposium, AD 220160 (Warfare Electrical Engineering Department, U.S. Army Research and Development Laboratories, Ft. Belvoir, Va., 1958), pp. 244–273."
  • Johnson's work was expressed in line pairs, but most people equate a line pair to a pair of pixels.
  • Johnson also looked at other goals such as determining orientation, but detection, recognition and identification are the most commonly-used today.

Thursday, October 3, 2013

Thoughts on the Future of the Microdisplay Business

If I were a shareholder of a microdisplay company such as eMagin or Kopin, I'd be a little worried about where future growth is going to come from.

For years, the microdisplay pitch was something like this: we make microdisplays for specialized applications - such as military products - where high performance are required, sometimes coupled with the ability to withstand harsh environments. One day, there will be a consumer market for such products in the form of virtual reality goggles or high-quantity of camera viewfinders, and this will allow us to reduce the price of our products and expand our reach. While this is coming, we make money by selling our specialized markets and doing contract research work.

This pitch is starting to look problematic. The consumer market is waking up, but not necessarily to the displays made by eMagin and Kopin.

In immersive virtual reality (e.g. not a see-through system), smartphone displays are a much more economical solution. Because more than a hundred million smartphones are sold every year, the cost of a high-resolution smartphone display can easily be less than 5% the cost of a comparable microdisplay. Microdisplay pricing has always been a chicken-and-an-egg game: prices can go down if quantities increase, but quantities will increase only if prices go down AND enough capital is available for production line and tooling investments.

Other technologies are also good candidates: pico projectors might become very popular for heads-up displays in cars and once large quantities will be made, they can also replace the microdisplay as a technology of choice.

Pico projectors are physically small which might be attractive to see-through goggles similar to Google Glass. The current generation of see-through consumer products does not seek to be high resolution nor wide field of view, and thus low-cost LCOS displays (Google is reportedly using Hynix) can provide a good solution for a high-brightness display that can be used outdoors. Karl Guttag had an interesting article on why Kopin's transmissive displays are not a good fit for these kind of applications.

One more thing on the subject of microdisplay prices. Though the financial reports do not reveal that microdisplays are a terrifically-profitable business, I suspect prices are also kept at some level because of "most favored nation" clauses to key customers such as perhaps the US government. Such clauses might force a microdisplay company that reduces prices to offer these reduced price levels to these 'most favored' customers. Thus if - for example - the US government is responsible for a large portion of a company's revenue and has a most-favored nation clause, any reduction in pricing beyond what is offered to the government will immediately results in significant loss of revenue once the US government prices are also reduced.

There will always be specialized applications where a display like eMagin's can be a perfect fit. Perhaps ones that requires very small physical size (such as when installed in a simulated weapon), or ones that can withstand extreme temperature and shock, or ones where quality is paramount and cost is secondary, but these do not sound like high-volume consumer applications.

The financial reports of both eMagin and Kopin reflect this reality. Both companies are currently losing money as they seek to address this reality.

What can be done to expand the business? One option is vertical integration. An opto-electronic system using a display needs additional components such as driver boards and optics beyond the display.  Today, these come from third-party vendors but one could imagine micro-display companies offering electronics and optics - or maybe even motion trackers - for small to medium-sized production runs. Another option which is currently pursued by Kopin is offering complete platforms and systems such as the Golden-i platform. Ostensibly, the margins on systems are much higher than the margin on individual components, especially as these become commodities. Over time, perhaps there is greater intellectual property there as well.

It will be interesting to see how this market shakes out in the upcoming months.

Full disclosure: I am not a shareholder of either company but my company uses eMagin microdisplays for several of our products.